Memorisation versus generalisation in pre-trained language models
File(s) Paper_Michael_Tanzer_ACL.pdf (1.07 MB)
Published version
Author(s)
Tänzer, Michael
Ruder, Sebastian
Rei, Marek
Type
Conference Paper
Abstract
State-of-the-art pre-trained language models
have been shown to memorise facts and per-
form well with limited amounts of training
data. To gain a better understanding of how
these models learn, we study their generali-
sation and memorisation capabilities in noisy
and low-resource scenarios. We find that the
training of these models is almost unaffected
by label noise and that it is possible to reach
near-optimal results even on extremely noisy
datasets. However, our experiments also show
that they mainly learn from high-frequency
patterns and largely fail when tested on low-
resource tasks such as few-shot learning and
rare entity recognition. To mitigate such lim-
itations, we propose an extension based on
prototypical networks that improves perfor-
mance in low-resource named entity recogni-
tion tasks.
have been shown to memorise facts and per-
form well with limited amounts of training
data. To gain a better understanding of how
these models learn, we study their generali-
sation and memorisation capabilities in noisy
and low-resource scenarios. We find that the
training of these models is almost unaffected
by label noise and that it is possible to reach
near-optimal results even on extremely noisy
datasets. However, our experiments also show
that they mainly learn from high-frequency
patterns and largely fail when tested on low-
resource tasks such as few-shot learning and
rare entity recognition. To mitigate such lim-
itations, we propose an extension based on
prototypical networks that improves perfor-
mance in low-resource named entity recogni-
tion tasks.
Date Acceptance
2022-02-24
Citation
1
Publisher
Association for Computational Linguistics
Volume
1
Copyright Statement
ACL materials are Copyright © 1963–2022 ACL; other materials are copyrighted by their respective copyright holders. Materials prior to 2016 here are licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 International License. Permission is granted to make copies for the purposes of teaching and research. Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License.
Identifier
https://aclanthology.org/2022.acl-long.521/
Source
The 60th Annual Meeting of the Association for Computational Linguistics (ACL 2022)
Publication Status
Published
Start Date
2022-05-22
Finish Date
2022-05-27
Coverage Spatial
Dublin, Ireland
