Public Opinion Quarterly

Comparing Speech-to-Text Algorithms for Transcribing Voice Data from Surveys

2025-10-02

With the increasing frequency of surveys being conducted via smartphones and tablets, the option of audio responses for providing open-ended responses has become more popular in research. This approach aims to improve the user experience of respondents and the quality of their answers. To circumvent the tedious task of transcribing each audio recording for analysis, previous studies have used the Google Cloud Speech-to-Text API to convert audio data to text. Extending previous research, we benchmark the Google Cloud service with state-of-the-art automatic speech recognition (ASR) systems from Meta (wav2vec 2.0), NVIDIA (NeMo), and OpenAI (Whisper). To do so, we use 100 randomly selected and recorded open-ended responses to popular social science survey questions. Additionally, we provide a basic, easy-to-understand introduction on how the Whisper ASR system works as well as code for implementation. By comparing Word Error Rates, we show that for our data the Google Cloud ASR service is outperformed by almost all other systems, highlighting the need to also consider other ASR systems.

Full text

DOI https://doi.org/10.1093/poq/nfaf056