Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
Michela Proietti, Roberto Capobianco, Mariya Toneva
ICML 2026 regular
Abstract (source: OpenReview · © authors)
Understanding the alignment between large language models (LLMs) and human brain activity can reveal computational principles underlying language processing. This work describes a pipeline to apply attribution methods to the brain-LLM alignment setting to identify the specific words most important for this alignment. As a case study, we leverage it to study a contentious research question about brain-LLM alignment: the relationship between brain alignment (BA) and next-word prediction (NWP). Across two naturalistic fMRI datasets, we find that BA and NWP rely on largely distinct word subsets: NWP exhibits recency and primacy biases with a focus on syntax, while BA prioritizes semantic and discourse-level information with a more targeted recency effect. This work advances our understanding of how LLMs relate to human language processing and highlights differences in feature reliance between BA and NWP. Beyond this study, our attribution method can be broadly applied to explore the cognitive relevance of model predictions in diverse language processing tasks.
Keywords
Metadata from BioTender-max/icml2026-ai-bio (CC0-1.0). Phở does not host any PDF; links point back to the source.
Related
Alignment between Brains and AI: Evidence for Convergent Evolution across Modalities, Scales and Training Trajectories
Guobin Shen, Dongcheng Zhao, Yiting Dong, Qian Zhang +1
Artificial and biological systems may converge on similar computational strategies despite different architectures and learning mechanisms—a form of convergent evolution. We test…
Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex
Abdulkadir Gokce, Yingtian Tang, Martin Schrimpf
Task-optimized neural networks are the leading in-silico models of sensory cortex, yet the field lacks a unified understanding of which modeling choices drive improved brain…
Learning Biophysical Models of Large-Scale Multineuronal Data To Enable Precise Neurostimulation
Amrith Lotlikar, Ian Christopher Tanoh, Praful K. Vasireddy, Andrew Lanpouthakoun +4
Multi-compartment Hodgkin–Huxley (HH) models provide a principled framework for predicting neural dynamics and responses to electrical stimulation. However, fitting HH biophysical…