Scientific Data

A Multi-view Dataset for Vietnamese Word-Level Sign Language Recognition

2026-08-14

This research introduces VSL400, a multi-view video dataset for isolated word-level recognition of Vietnamese Sign Language (VSL). The dataset contains 74,259 manually annotated video clips covering 400 glosses performed by 28 signers, including deaf student signers and additional trained VSL signers. Each signing instance was recorded simultaneously from three synchronized RGB camera views: front, left, and right. This design captures complementary visual perspectives of sign articulation and supports single-view, cross-view, and multi-view recognition studies. To facilitate reproducible research, we provide a preprocessing pipeline for temporal boundary detection, segmentation, and spatial normalization, together with structured metadata and reference code. Metadata, documentation, and processing code are available from Zenodo, while the de-identified human video files are shared through controlled access governed by a Data Usage Agreement. VSL400 addresses a major resource gap for Vietnamese Sign Language and provides a practical reference dataset for research on isolated sign recognition in a low-resource language setting.

Full text

DOI https://doi.org/10.1038/s41597-026-08040-2