Linguistic diversity in machine learning training data: Content analysis of English and Finnish soundscape descriptions in audio captioning
Hekanaho, Laura; Tuuri, Emilia; Surakka, Maija; Hirvonen, Maija (2026-06-11)
Hekanaho, Laura
Tuuri, Emilia
Surakka, Maija
Hirvonen, Maija
11.06.2026
PLoS ONE
e0350043
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202606157407
https://urn.fi/URN:NBN:fi:tuni-202606157407
Kuvaus
Peer reviewed
Tiivistelmä
This article examines descriptions of auditory experiences used as training data in machine learning. While such training datasets play a key role in machine learning, only a few studies have investigated these training datasets from a linguistic perspective. Making an important contribution to the English-focused field, we compare descriptions produced in two typologically distant languages: English and Finnish. The English descriptions were produced by sighted participants, and the Finnish data both by sighted and non-sighted individuals, offering insight into variation in language use based on sensory experience. First, we provide an overview of the three datasets with a descriptive corpus analysis. The corpus analysis reveals greater variation among the Finnish corpora, in comparison with the English corpus. We attribute this to the use of L2 speakers as informants for the English data, considering potential data quality issues from a linguistic perspective. Second, we report on a qualitative content analysis conducted on a smaller sample of 150 descriptions, examining which elements the participants had identified and chosen to verbalize. In particular, the analysis indicates that the visually disabled participants make use of slightly richer lexical resources, especially in terms of verbalising spatial structure. In addition, a methodological finding of our study captures how the corpus-level and the discourse-level analyses lead to distinct, even seemingly contradictory results: while the corpus analysis focuses on token distributions, the qualitative discourse-level analysis reveals which different semantic domains are prevalent in different subcorpora. This article contributes to the growing research on cross-linguistic and language-internal diversity in the field of machine perception.
Kokoelmat
- TUNICRIS-julkaisut [24977]
