The Dialogue to Deployment Pipeline: A Multivocal Literature Review
Rizwani, Abdul Wasay (2026)
Rizwani, Abdul Wasay
2026
Master's Programme in Computing Sciences and Electrical Engineering
Informaatioteknologian ja viestinnän tiedekunta - Faculty of Information Technology and Communication Sciences
This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.
Hyväksymispäivämäärä
2026-06-16
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202606167527
https://urn.fi/URN:NBN:fi:tuni-202606167527
Tiivistelmä
Modern software development suffers a persistent translation loss: client intent expressed in natural language must pass through multiple human intermediaries before it reaches running code. This thesis investigates how three maturing AI technologies (automatic speech recognition, large language models, and multi-agent systems) can be composed into a continuous pipeline that converts unstructured stakeholder conversations directly into deployable Minimum Viable Products (MVPs).
The work is structured as a Multivocal Literature Review (MLR), following Garousi, Felderer, and M¨antyl¨a’s guidelines for including grey literature alongside peer-reviewed sources, with PRISMA 2020 used as the reporting scaffold. A corpus of 45 peer-reviewed primary studies and 18 grey-literature sources, retrieved from IEEE Xplore, ACM Digital Library, Scopus, Springer-Link, and arXiv over the period 2023–2026, is synthesised around three research questions aligned with three pipeline layers: the Listener layer (RQ1: speech-to-text and requirement extraction), the Developer layer (RQ2: multi-agent code generation and MVP assembly), and the Reliability layer (RQ3: hallucination mitigation and long-context retention).
Key findings are as follows. On RQ1, state-of-the-art STT systems achieve word error rates below 3% on clean audio, but degrade significantly on overlapping speech present in 18–24% of agile refinement meetings; LLM-based multi-agent extraction reliably recovers functional requirements yet systematically under-recalls non-functional ones. On RQ2, academic frameworks (MetaGPT, ChatDev, SWE-Agent) demonstrate strong performance on synthetic benchmarks (HumanEval, SWE-bench) but produce incomplete cloud-native deployment artefacts; commercial vibe-coding platforms close the deployment gap but lack rigorous verification. On RQ3, Chain of Agents outperforms vanilla RAG on long-context retention; Knowledge Graph grounding (GraphRAG, Think-on-Graph) most reliably reduces cross-document hallucination; multi-agent debate provides a low-cost complementary signal.
These findings are integrated into the Golden Path: a synthesised three-layer reference architecture combining Whisper-based transcription with speaker diarisation, Chain-of-Agents aggregation, multi-agent requirements drafting grounded against a persistent project knowledge graph, SWE-Agent-style programmer agents with automated verification, and a dedicated De-vOps agent producing complete cloud-native artefacts. The architecture addresses identified failure modes but has not yet been instantiated end-to-end; empirical validation is identified as the primary direction for future work.
The work is structured as a Multivocal Literature Review (MLR), following Garousi, Felderer, and M¨antyl¨a’s guidelines for including grey literature alongside peer-reviewed sources, with PRISMA 2020 used as the reporting scaffold. A corpus of 45 peer-reviewed primary studies and 18 grey-literature sources, retrieved from IEEE Xplore, ACM Digital Library, Scopus, Springer-Link, and arXiv over the period 2023–2026, is synthesised around three research questions aligned with three pipeline layers: the Listener layer (RQ1: speech-to-text and requirement extraction), the Developer layer (RQ2: multi-agent code generation and MVP assembly), and the Reliability layer (RQ3: hallucination mitigation and long-context retention).
Key findings are as follows. On RQ1, state-of-the-art STT systems achieve word error rates below 3% on clean audio, but degrade significantly on overlapping speech present in 18–24% of agile refinement meetings; LLM-based multi-agent extraction reliably recovers functional requirements yet systematically under-recalls non-functional ones. On RQ2, academic frameworks (MetaGPT, ChatDev, SWE-Agent) demonstrate strong performance on synthetic benchmarks (HumanEval, SWE-bench) but produce incomplete cloud-native deployment artefacts; commercial vibe-coding platforms close the deployment gap but lack rigorous verification. On RQ3, Chain of Agents outperforms vanilla RAG on long-context retention; Knowledge Graph grounding (GraphRAG, Think-on-Graph) most reliably reduces cross-document hallucination; multi-agent debate provides a low-cost complementary signal.
These findings are integrated into the Golden Path: a synthesised three-layer reference architecture combining Whisper-based transcription with speaker diarisation, Chain-of-Agents aggregation, multi-agent requirements drafting grounded against a persistent project knowledge graph, SWE-Agent-style programmer agents with automated verification, and a dedicated De-vOps agent producing complete cloud-native artefacts. The architecture addresses identified failure modes but has not yet been instantiated end-to-end; empirical validation is identified as the primary direction for future work.