Sensory context improves language prediction in humans and LLMs
2026-08-26
Language is a fundamental human capacity. Large language models (LLMs) have presented the first viable model of language outside of humans, yet how these models learn and use language differs significantly from humans. Here, we compare LLMs and humans predicting language with varying levels of sensory information—from disembodied written text to audiovisual videos of speakers—to demonstrate that, in both humans and LLMs, sensory context is critical for optimal performance. We asked human participants to predict upcoming words within narratives presented as either audiovisual, audio-only, or written language ( N = 1 , 500 total; 500 per modality), focusing on content words that convey essential meaning. We compared these predictions across modalities as well as to predictions generated by LLMs. Human predictions were overall more accurate than LLMs, regardless of the sensory modality of presentation. Compared to written language, both audiovisual and audio-only language increased the accuracy and consensus of human predictions and decreased alignment with LLMs. We identified that prosody, a signal that conveys important linguistic information via the auditory channel, partially drove both the observed advantage in accuracy within humans and divergence from LLMs. Integrating multimodal information—i.e., prosody, auditory, or audiovisual information—with the representation of language learned during LLM training improved models’ next-word prediction performance and increased the efficiency of language learning. These findings demonstrate that sensory contexts are foundational to human-like language behavior, and that these contexts can enrich and accelerate language acquisition within LLMs similar to what is observed within human development.