The impact of token manipulation in Natural Language Generation systems
File(s)
Author(s)
Vedd, Nihir
Type
Thesis
Abstract
In recent years, Natural Language Generation (NLG) systems have achieved human-level performance on a variety of tasks, primarily using encoder-decoder or decoder-only architectures. Training these decoders on downstream data has been enhanced by techniques that manipulate token presence. For instance, BART employs methods such as token masking, sentence permutation, and token deletion to corrupt inputs. We introduce the term "token manipulation" to describe these methods that alter token presence to impact language models. Our hypothesis posits that token manipulation can improve NLG system performance, which we classify into input-side and output-side approaches. Input-side manipulation occurs before a model processes inputs, whereas output-side manipulation happens after.
Our research includes novel token manipulation strategies for both uni-modal and multi-modal tasks. Initially, we explored masking during fine-tuning, noting substantial improvements over baseline models. Encouraged by input-side manipulation’s success, we tested output-side manipulation by applying artificial tokenisations to an auxiliary decoder, which also yielded significant enhancements without increasing parameter count or computational complexity.
Furthermore, we applied input-side token manipulation to a multi-modal Visual Question Generation (VQG) task, introducing three new strategies for token imputation during modeling, achieving significant performance gains.
Combining our input-side and output-side models resulted in our most effective model, showing a 14.5% increase in metrics across various uni-lingual generative tasks. In translation benchmarks, we observed a 2.4 BLEU-4 improvement in IWSLT DE-EN, 0.8 in WMT '16 DE-EN, and 1.6 in WMT '16 EN-RO. Our VQG model surpassed state-of-the-art results by more than 9 BLEU-4 points at publication.
Our research includes novel token manipulation strategies for both uni-modal and multi-modal tasks. Initially, we explored masking during fine-tuning, noting substantial improvements over baseline models. Encouraged by input-side manipulation’s success, we tested output-side manipulation by applying artificial tokenisations to an auxiliary decoder, which also yielded significant enhancements without increasing parameter count or computational complexity.
Furthermore, we applied input-side token manipulation to a multi-modal Visual Question Generation (VQG) task, introducing three new strategies for token imputation during modeling, achieving significant performance gains.
Combining our input-side and output-side models resulted in our most effective model, showing a 14.5% increase in metrics across various uni-lingual generative tasks. In translation benchmarks, we observed a 2.4 BLEU-4 improvement in IWSLT DE-EN, 0.8 in WMT '16 DE-EN, and 1.6 in WMT '16 EN-RO. Our VQG model surpassed state-of-the-art results by more than 9 BLEU-4 points at publication.
Version
Open Access
Date Issued
2024-03
Date Awarded
2024-08
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Specia, Lucia
Rei, Marek
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)