Multi-Agent Topic Modeling : An Experimental Study Using Synthetic Data
Kylliäinen, Rolle (2026)
Kylliäinen, Rolle
2026
Tieto- ja sähkötekniikan kandidaattiohjelma - Bachelor's Programme in Computing and Electrical Engineering
Informaatioteknologian ja viestinnän tiedekunta - Faculty of Information Technology and Communication Sciences
This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.
Hyväksymispäivämäärä
2026-06-17
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202606157487
https://urn.fi/URN:NBN:fi:tuni-202606157487
Tiivistelmä
Topic modeling approaches such as LDA, BERTopic and K-Means frequently produce inconsistent and varying results when applied to short, domain specific texts, making it difficult to obtain a stable and interpretable topic representations. This thesis designs, implements, and evaluates a multi-agent framework that coordinates these three structurally different topic modeling approaches through a four-agent workflow, reconciling their outputs into a unified topic representation. The framework is evaluated on a synthetic dataset of 1,000 safety notification documents spanning 15 predefined thematic categories, generated by prompting a large language model with samples drawn from a real industrial dataset.
Experimental results provide preliminary evidence that the consensus mechanism achieves the highest C_v coherence of 0.726 ± 0.007 across five benchmark runs, outperforming the strongest individual baseline BERTopic (0.690 ± 0.014). On LLM-based evaluation, BERTopic scored marginally higher than the consensus, reflecting a trade-off between thematic coverage and topic distinctness introduced by the pipeline. The two evaluation methods produce consistent model rankings but differ in sensitivity, with the LLM evaluator capturing semantic dimensions such as interpretability that co-occurrence statistics do not reflect.
The primary contribution of this thesis is the design and preliminary evaluation of a multi-agent topic modeling system that demonstrates how structurally different topic modeling approaches can be coordinated through quality-weighted consensus to produce more stable and coherent topic representations than any individual model would achieve. The results also provide preliminary evidence that LLM-based evaluation may capture aspects of topic quality that traditional coherence metrics miss, supporting its use as a complementary evaluation instrument in topic modeling research.
Experimental results provide preliminary evidence that the consensus mechanism achieves the highest C_v coherence of 0.726 ± 0.007 across five benchmark runs, outperforming the strongest individual baseline BERTopic (0.690 ± 0.014). On LLM-based evaluation, BERTopic scored marginally higher than the consensus, reflecting a trade-off between thematic coverage and topic distinctness introduced by the pipeline. The two evaluation methods produce consistent model rankings but differ in sensitivity, with the LLM evaluator capturing semantic dimensions such as interpretability that co-occurrence statistics do not reflect.
The primary contribution of this thesis is the design and preliminary evaluation of a multi-agent topic modeling system that demonstrates how structurally different topic modeling approaches can be coordinated through quality-weighted consensus to produce more stable and coherent topic representations than any individual model would achieve. The results also provide preliminary evidence that LLM-based evaluation may capture aspects of topic quality that traditional coherence metrics miss, supporting its use as a complementary evaluation instrument in topic modeling research.
Kokoelmat
- Kandidaatintutkielmat [11870]
