Hyppää sisältöön
    • Suomeksi
    • In English
Trepo
  • Suomeksi
  • In English
  • Kirjaudu
Näytä viite 
  •   Etusivu
  • Trepo
  • TUNICRIS-julkaisut
  • Näytä viite
  •   Etusivu
  • Trepo
  • TUNICRIS-julkaisut
  • Näytä viite
JavaScript is disabled for your browser. Some features of this site may not work without it.

FinnAffect: An Affective Speech Corpus for Spontaneous Finnish

Lahtinen, Kalle; Mustanoja, Liisa; Räsänen, Okko (2025-11)

 
Avaa tiedosto
Finnaffect_Lahtinen_ym.pdf (1.410Mt)
Lataukset: 



Lahtinen, Kalle
Mustanoja, Liisa
Räsänen, Okko
11 / 2025

Speech Communication
103327
doi:10.1016/j.specom.2025.103327
Näytä kaikki kuvailutiedot
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-2025112610941

Kuvaus

Peer reviewed
Tiivistelmä
Affective expression plays a major role in everyday spoken and written language. In order to study how affect is expressed by Finnish language users in day-to-day life, data consisting of samples from naturalistic and unscripted contexts is required. The present work describes the first spontaneous speech corpus for Finnish with affect-related annotations, containing 12,000 transcribed samples of unscripted speech paired with continuous-valued scores of valence and arousal marked by five native Finnish speakers. We first describe the creation of the corpus, based on combining speech samples from three large-scale Finnish speech corpora, from which we chose samples for annotation using an active learning-based affect mining approach. We then report characteristics of the resulting corpus and annotation consistency, followed by speech emotion recognition (SER) experiments with several classifiers and regression models to test the feasibility of the corpus for SER system development and evaluation. Annotation analyses reveal mean Pearson correlations between annotator scores and the mean of all annotators to be rho = 0.856 for valence and rho = 0.898 for arousal. The SER experiments on discretized labels result in an average unweighted average recall (UAR) of 0.458 for ternary valence classification and 0.719 for binary arousal classification using a fine-tuned ExHuBERT model for valence prediction and a support vector machine (SVM) classifier for arousal prediction, reaching comparable levels to those reported earlier for spontaneous speech. For the regression task, concordance correlation coefficients of 0.270 and 0.689 were obtained for valence and arousal, respectively, when using a WavLM-based model trained on MSP-Podcast corpus and fine-tuned on the target data. Overall, the analyses suggest that the corpus provides a feasible basis for later study on affective expression in spontaneous Finnish.
Kokoelmat
  • TUNICRIS-julkaisut [25334]
Kalevantie 5
PL 617
33014 Tampereen yliopisto
oa[@]tuni.fi | Tietosuoja | Saavutettavuusseloste
 

 

Selaa kokoelmaa

TekijätNimekkeetTiedekunta (2019 -)Tiedekunta (- 2018)Tutkinto-ohjelmat ja opintosuunnatAvainsanatJulkaisuajatKokoelmat

Omat tiedot

Kirjaudu sisäänRekisteröidy
Kalevantie 5
PL 617
33014 Tampereen yliopisto
oa[@]tuni.fi | Tietosuoja | Saavutettavuusseloste