The Tangrams Codeswitching Corpus
- Anne L. Beatty-Martínez
- Jiaze Li
- Yumeng Shen
- Stanislav Mulík
- Robert D. Hawkins
- Rosa E. Guzzardo Tamargo
- Paola E. Dussias
2026-08-19
Most research in the language sciences has centered on monolingual data, limiting the relevance of its insights for multilingual populations. The Tangrams Codeswitching Corpus , a Spanish-English dataset containing 30,858 words from 45 dyadic repeated reference games, bridges this gap by offering a richly annotated dataset on bilingual communication and codeswitching. In addition to the corpus, the dataset includes detailed metadata on language experience and proficiency of 90 habitual codeswitchers, enabling studies of individual differences in social coordination and switching behavior. This publicly available resource is valuable for exploring bilinguals’ language choices and the interactional dynamics through which they communicate, collaborate, and coordinate with one another. By providing a transdisciplinary perspective on codeswitching, this dataset offers a unique opportunity to explore how bilinguals adapt their language choices and communicative behavior in everyday interaction.