Jacobian Scopes: token-level causal attributions in LLMs
File(s) 2601.16407v2.pdf (7.52 MB)
Preprint version
OA Location
Author(s)
Liu, Toni JB
Zadeoğlu, Baran
Boulle, Nicolas
Sarfati, Raphael
Earls, Christopher J
Type
preprint
Abstract
Large language models (LLMs) make nexttoken predictions based on clues present in their context, such as semantic descriptions and incontext examples. Yet, elucidating which prior tokens most strongly influence a given prediction remains challenging due to the proliferation of layers and attention heads in modern architectures. We propose Jacobian Scopes, a suite of gradient-based, token-level causal attribution methods for interpreting LLM predictions. Grounded in perturbation theory and information geometry, Jacobian Scopes quantify how input tokens influence various aspects of a model's prediction, such as specific logits, the full predictive distribution, and model uncertainty (effective temperature). Through case studies spanning instruction understanding, translation, and in-context learning (ICL), we demonstrate how Jacobian Scopes reveal implicit political biases, uncover word-and phrase-level translation strategies, and shed light on recently debated mechanisms underlying in-context time-series forecasting. To facilitate exploration of Jacobian Scopes on custom text, we open-source our implementations and provide a cloud-hosted interactive demo at https://huggingface.co/spaces/ Typony/JacobianScopes.
Date Issued
2026-01-23
Citation
arXiv, 2026
Journal / Book Title
arXiv
Copyright Statement
Copyright © 2026 The Authors.
Description
Preprint version
