Scientific Data

A Figure-Text Pair Dataset from Map-related Literature for Multimodal Model Research

2026-09-05

High-quality multimodal datasets are essential for developing vision-language models, yet publicly available figure-text resources in specialized scientific domains remain limited. To address this gap, we present a large-scale figure-text pair dataset constructed from map-related scientific literature (FTPD-ML). The dataset was derived from 96,859 scientific publications in cartography, geography, remote sensing, and related disciplines, and provides 75,702 publicly redistributable figure-text pairs under Creative Commons Attribution (CC BY) licenses with multiple levels of textual representations, including original figure captions, standardized caption variants, and contextual paragraph descriptions, together with bibliographic metadata. To ensure compliance with copyright and licensing requirements, records from non-redistributable publications are represented by metadata only. The dataset supports a range of multimodal research tasks, including image-text retrieval, image captioning, scientific document understanding, and cartography-oriented vision-language studies.

Full text

DOI https://doi.org/10.1038/s41597-026-08248-2