Hyppää sisältöön
    • Suomeksi
    • In English
Trepo
  • Suomeksi
  • In English
  • Kirjaudu
Näytä viite 
  •   Etusivu
  • Trepo
  • TUNICRIS-julkaisut
  • Näytä viite
  •   Etusivu
  • Trepo
  • TUNICRIS-julkaisut
  • Näytä viite
JavaScript is disabled for your browser. Some features of this site may not work without it.

Leveraging Large Language Models for Twitch Video Clip Understanding and Summarization

Lindroos, Jari; Koskimaa, Raine; Peltonen, Jaakko; Välisalo, Tanja; Toivanen, Ida (2026-03-22)

 
Avaa tiedosto
Leveraging_Large_Language_Models_for_Twitch_Video_Clip_Understanding_and_Summarization.pdf (1.064Mt)
Lataukset: 



Lindroos, Jari
Koskimaa, Raine
Peltonen, Jaakko
Välisalo, Tanja
Toivanen, Ida
22.03.2026

1
doi:10.1145/3789624.3789641
Näytä kaikki kuvailutiedot
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202606177571

Kuvaus

Peer reviewed
Tiivistelmä
Livestreaming platforms like Twitch generate a vast amount of multimodal data, including the synchronized video, audio, and the chat from the audience. Analyzing and summarizing this information is difficult at scale. This paper presents an exploratory study which introduces and evaluates an end-to-end framework that leverages Multimodal Large Language Models (MLLMs) from the Google Gemini family to automatically understand and summarize short-form Twitch video clips. Our methodology uses a Chain-of-Thought-based prompt to identify the key audio-visual moments and the reactions of the chat. This information is then correlated across the modalities to provide an overall comprehensive summary. The quality of these summaries were assessed through a human-centric evaluation on a curated dataset of 56 clips. Our analysis suggests a clear generational leap in performance, with the Gemini 2.5 models significantly outperforming the previous series. This improvement represents a qualitative shift from descriptive to analytical reasoning, including the ability to understand domain-specific in-game context. We also observe, through the modality ablation study, a fundamental improvement in inherent video understanding capabilities when solely utilizing audio-visual information as input. This work serves as foundational research for the application of utilizing MLLMs to livestreaming content, providing a groundwork for future studies in Long-form Video Understanding, such as the analysis of full matches or entire tournament broadcasts in the context of esports.
Kokoelmat
  • TUNICRIS-julkaisut [25372]
Kalevantie 5
PL 617
33014 Tampereen yliopisto
oa[@]tuni.fi | Tietosuoja | Saavutettavuusseloste
 

 

Selaa kokoelmaa

TekijätNimekkeetTiedekunta (2019 -)Tiedekunta (- 2018)Tutkinto-ohjelmat ja opintosuunnatAvainsanatJulkaisuajatKokoelmat

Omat tiedot

Kirjaudu sisäänRekisteröidy
Kalevantie 5
PL 617
33014 Tampereen yliopisto
oa[@]tuni.fi | Tietosuoja | Saavutettavuusseloste