Advancing large language models for comprehensive scientific assistance
File(s)
Author(s)
Korkmaz, Buse Sibel
Type
Thesis
Abstract
Despite their transformative potential, large language models remain ill-suited for scientific inquiry. They hallucinate citations, miss quantitative evidence in tables, construct arguments without attribution, and amplify biases. These failures undermine trust where researchers need it most. This thesis reimagines large language models as accountable companions that augment scientific judgment through grounding, argumentation, and responsible generation.
We develop capability-specific enhancements targeting three critical bottlenecks. First, we demonstrate that existing models fundamentally lack scientific table understanding, a capability essential for evidence-based reasoning. Through curriculum training that progresses from general table reasoning to domain-specific scientific tables, we show that models which truly comprehend tabular evidence, rather than superficially processing table captions, develop representations that substantially improve peer review score prediction. This finding establishes that deep engagement with quantitative data is necessary for models to assess scientific claims as researchers do.
Second, we treat scientific argumentation as fundamentally distinct from generic reasoning. Through a pipeline combining reasoning-aware distillation, multi-task learning, and reinforcement learning with composite rewards, we teach models to locate claims, attribute evidence, and draft targeted critiques with measurable gains across argument mining, generation, and discourse evaluation. Third, we address the alignment tax, the trade-off between safety and performance, by introducing methods that preserve factual faithfulness while reducing toxicity and bias without degrading language quality.
These capabilities converge in a retrieval-augmented workflow designed for transparency. Outputs expose sources, reveal reasoning, and acknowledge limitations. Ablation studies confirm that modules contribute additively and interact constructively, delivering grounded, auditable, and bias-aware assistance. The collective findings reframe large language models from error-prone editorial aids to trustworthy instruments for evidence-based inquiry that scientists can genuinely rely on.
We develop capability-specific enhancements targeting three critical bottlenecks. First, we demonstrate that existing models fundamentally lack scientific table understanding, a capability essential for evidence-based reasoning. Through curriculum training that progresses from general table reasoning to domain-specific scientific tables, we show that models which truly comprehend tabular evidence, rather than superficially processing table captions, develop representations that substantially improve peer review score prediction. This finding establishes that deep engagement with quantitative data is necessary for models to assess scientific claims as researchers do.
Second, we treat scientific argumentation as fundamentally distinct from generic reasoning. Through a pipeline combining reasoning-aware distillation, multi-task learning, and reinforcement learning with composite rewards, we teach models to locate claims, attribute evidence, and draft targeted critiques with measurable gains across argument mining, generation, and discourse evaluation. Third, we address the alignment tax, the trade-off between safety and performance, by introducing methods that preserve factual faithfulness while reducing toxicity and bias without degrading language quality.
These capabilities converge in a retrieval-augmented workflow designed for transparency. Outputs expose sources, reveal reasoning, and acknowledge limitations. Ablation studies confirm that modules contribute additively and interact constructively, delivering grounded, auditable, and bias-aware assistance. The collective findings reframe large language models from error-prone editorial aids to trustworthy instruments for evidence-based inquiry that scientists can genuinely rely on.
Version
Open Access
Date Issued
2025-11-05
Date Awarded
2026-02-01
Copyright Statement
Attribution-NonCommercial-ShareAlike 4.0 International Licence (CC BY NC-SA)
Advisor
del Rio Chanona, Antonio
Publisher Department
Department of Chemical Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
