Hyppää sisältöön
    • Suomeksi
    • In English
Trepo
  • Suomeksi
  • In English
  • Kirjaudu
Näytä viite 
  •   Etusivu
  • Trepo
  • Väitöskirjat
  • Näytä viite
  •   Etusivu
  • Trepo
  • Väitöskirjat
  • Näytä viite
JavaScript is disabled for your browser. Some features of this site may not work without it.

Neural Speech Separation in Real-world Conversational Scenarios

Wang, Yuzhu (2026)

 
Avaa tiedosto
978-952-03-4754-3.pdf (25.41Mt)
Lataukset: 



Wang, Yuzhu
Tampere University
2026

Tieto- ja sähkötekniikan tohtoriohjelma - Doctoral Programme in Computing and Electrical Engineering
Informaatioteknologian ja viestinnän tiedekunta - Faculty of Information Technology and Communication Sciences
This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.
Väitöspäivä
2026-10-06
Näytä kaikki kuvailutiedot
Julkaisun pysyvä osoite on
https://urn.fi/URN:ISBN:978-952-03-4754-3
Tiivistelmä
The human auditory system can perceptually segregate overlapping sounds into distinct auditory streams while perceiving their spatial locations and can actively focus on a single speaker among multiple concurrent speakers, the latter known as the cocktail party effect. Inspired by these perceptual abilities, speech separation aims to develop computational methods that estimate individual source signals from mixtures, with applications spanning automatic meeting transcription, hearing aids, and voice-controlled assistants. Neural methods have driven substantial progress in the field, yet these advances have largely come at the cost of simplifying assumptions that rarely hold in real conversational scenarios. Most existing methods require the number of speakers to be fixed and known, operate on short recordings, and assume that sources remain stationary. However, conversations involve a varying number of participants, extend over arbitrary durations with each speaker producing multiple utterances separated by silence, and feature speakers who move throughout the interaction. This thesis belongs to the fields of signal processing and machine learning, and investigates neural approaches to speech separation in real-world conversational scenarios.

Three challenges are addressed in this thesis. For the unknown speaker number problem, an attractor-based approach is proposed that performs speaker counting, speaker activity detection, and separation within a joint architecture, allowing the number of output streams to be determined from the input. For arbitrary-length recordings, two approaches are developed. The first employs full-band and sub-band modeling to enable direct inference on recordings substantially longer than the training segments with-out segmentation. The second is a dynamic clustering method that re-solves cross-segment permutation ambiguity using quality-weighted reference embeddings. For moving sources, the thesis investigates attention-based spatial filtering for adaptive tracking of time-varying source positions, and proposes a dual-branch architecture that processes spectral and spatial features through parallel branches, since these two feature types evolve at fundamentally different temporal scales.

The proposed methods are experimentally validated across diverse acoustic conditions, demonstrating the ability to handle unknown speaker numbers, process recordings substantially longer than the training segments, and separate moving sources with time-varying spatial characteristics. In particular, the observation that spectral and spatial features evolve at fundamentally different temporal scales suggests that system designs respecting such multi-scale characteristics may offer a broadly applicable principle for speech and audio processing beyond the specific challenges addressed in this thesis.
Kokoelmat
  • Väitöskirjat [5374]
Kalevantie 5
PL 617
33014 Tampereen yliopisto
oa[@]tuni.fi | Tietosuoja | Saavutettavuusseloste
 

 

Selaa kokoelmaa

TekijätNimekkeetTiedekunta (2019 -)Tiedekunta (- 2018)Tutkinto-ohjelmat ja opintosuunnatAvainsanatJulkaisuajatKokoelmat

Omat tiedot

Kirjaudu sisäänRekisteröidy
Kalevantie 5
PL 617
33014 Tampereen yliopisto
oa[@]tuni.fi | Tietosuoja | Saavutettavuusseloste