Sociological Methodology

Joint Text-and-Image Clustering for Social Science Research

2025-11-24

Automated text analysis is becoming extremely popular and image analysis is gaining interest. However, multimodal analysis that combines both text and image information remains rare, even though many real-world data are intrinsically multimodal, such as social media posts. The authors compare three practical workflows for clustering text–image pairs: (1) label-level combination, which clusters text and image separately and combines the resulting labels; (2) vector-level combination, which clusters concatenated embeddings extracted from each modality; and (3) joint embedding, which clusters unified representations from multimodal embedding models such as Contrastive Language-Image Pre-training. The authors also introduce a set of reusable evaluation tools to help researchers compare, validate, and benchmark multimodal clustering workflows: adjusted mutual information to assess text–image alignment, the S_DbW index to evaluate number of clusters, and within-cluster consistency to validate interpretability. The authors validate the methods on a Chinese protest data set from social media with 336,921 text–image pairs and test robustness and scope conditions using a smaller U.S. news data set on gun violence with 1,297 news headlines. The authors find that when text and image provide distinct, nonoverlapping information, the second and third methods outperform the first. This study serves as a bridge between the text-as-data and image-as-data communities.

Full text

DOI https://doi.org/10.1177/00811750251382929