Generative language model compression with algebraic approaches
File(s)
Author(s)
Xu, Mingxue
Type
Thesis
Abstract
Generative language models have empowered daily and professional AI applications, yet their internal mechanisms and deployment on affordable hardware remain underexplored. This thesis addresses generative language model compression, which serves dual purposes: understanding the inner workings of language models and enabling lightweight deployment.
The work targets three gaps: (1) current methods rely on empirical heuristics rather than principled theories, lacking clear connections between compression parameters and fundamental properties such as capacity and expressivity; (2) compression pipelines are bespoke scripts tightly coupled to specific cases, hindering reproducibility and systematic comparison; (3) many techniques assume abundant retraining resources, poorly matching training-free adaptation on lower-end devices.
To address (1), the thesis develops a geometric and algebraic framework that treats embeddings, attention, and feed-forward layers as compositions of linear operators on structured subspaces. This perspective reveals where low-rank structure is compatible with model semantics, unifying matrix and tensor factorisations under one taxonomy grounded in direct products, tensor products, and contractions. To address (2), the thesis introduces an object-oriented compression toolbox that operationalises these ideas through reusable abstractions, separating model management, algorithms, evaluation, and profiling behind stable interfaces for systematic comparison across open-source models. To address (3), a runtime study of tensor-train decomposition for embedding layers, which consume nearly half the parameters in sub-billion parameter models, on Raspberry Pi 5 shows that compression preserves or slightly improves accuracy while halving energy usage with negligible latency overhead.
Together, this thesis bridges theory, software engineering, and hardware-aware evaluation, illuminating language models through geometric subspaces and linear operators for controllable compression with predictable trade-offs.
The work targets three gaps: (1) current methods rely on empirical heuristics rather than principled theories, lacking clear connections between compression parameters and fundamental properties such as capacity and expressivity; (2) compression pipelines are bespoke scripts tightly coupled to specific cases, hindering reproducibility and systematic comparison; (3) many techniques assume abundant retraining resources, poorly matching training-free adaptation on lower-end devices.
To address (1), the thesis develops a geometric and algebraic framework that treats embeddings, attention, and feed-forward layers as compositions of linear operators on structured subspaces. This perspective reveals where low-rank structure is compatible with model semantics, unifying matrix and tensor factorisations under one taxonomy grounded in direct products, tensor products, and contractions. To address (2), the thesis introduces an object-oriented compression toolbox that operationalises these ideas through reusable abstractions, separating model management, algorithms, evaluation, and profiling behind stable interfaces for systematic comparison across open-source models. To address (3), a runtime study of tensor-train decomposition for embedding layers, which consume nearly half the parameters in sub-billion parameter models, on Raspberry Pi 5 shows that compression preserves or slightly improves accuracy while halving energy usage with negligible latency overhead.
Together, this thesis bridges theory, software engineering, and hardware-aware evaluation, illuminating language models through geometric subspaces and linear operators for controllable compression with predictable trade-offs.
Version
Open Access
Date Issued
2025-12-31
Date Awarded
2026-05-01
Copyright Statement
Attribution-NonCommercial 4.0 International Licence (CC BY-NC)
License URL
Advisor
Mandic, Danilo
Publisher Department
Department of Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
