Hyppää sisältöön
    • Suomeksi
    • In English
Trepo
  • Suomeksi
  • In English
  • Kirjaudu
Näytä viite 
  •   Etusivu
  • Trepo
  • Opinnäytteet - ylempi korkeakoulututkinto
  • Näytä viite
  •   Etusivu
  • Trepo
  • Opinnäytteet - ylempi korkeakoulututkinto
  • Näytä viite
JavaScript is disabled for your browser. Some features of this site may not work without it.

Agentic Data Visualization for GenAI Assistants: Adapting LIDA as an MCP Server

Liukko, Väinö (2026)

 
Avaa tiedosto
LiukkoVaino.pdf (582.7Kt)
Lataukset: 



Liukko, Väinö
2026

Tietotekniikan DI-ohjelma - Master's Programme in Information Technology
Informaatioteknologian ja viestinnän tiedekunta - Faculty of Information Technology and Communication Sciences
Hyväksymispäivämäärä
2026-07-31
Näytä kaikki kuvailutiedot
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202607318673
Tiivistelmä
This thesis began as an opportunity to explore new multimodal features for FunctionAI, Solita's closed-source internal GenAI assistant. Automated visualization generation was chosen as the focus of this study. A literature review identified LIDA, an LLM-based grammar-agnostic visualization library, as the most promising option. The review also found that most academic visualization tools had restrictive licenses or were not maintained, making enterprise adoption difficult; later in the thesis, it became clear that even LIDA required significant adaptation to be usable.

The work followed the Design Science Research methodology across two development iterations. In the first phase, LIDA was integrated directly into FunctionAI as a plugin. This revealed that executing LLM-generated Python code exposes the host to arbitrary code execution through prompt injection and that the target system lacked a well-defined plugin API. These findings motivated a second phase in which the artifact was redesigned as a stateless MCP server. In this design, the MCP sampling primitive delegates model selection to the client, making the server LLM-agnostic and enabling agentic usage patterns, while placing code execution in Azure Functions reduces the security risk of running LLM-generated code.

The sandboxing results were mixed. Azure provides effective network-level, filesystem, and timeout isolation. However, Azure Functions does not provide process-level isolation for generated-code execution: LLM-generated code can modify shared Python library objects, and such modifications persist across subsequent requests within the same worker process, making process-level state reset necessary. Improving this isolation is left for future work.

The LIDA MCP server was evaluated using the visualization error rate (VER) benchmark from the original LIDA paper, replicated on 59 vega-datasets across multiple plotting libraries and two test runs: an initial run with GPT-5.4-nano, Llama 4 Scout, and Gemini 2.5 Flash Lite, and a second run with GPT-5.4-nano, GPT-5.4-mini, and GPT-5.3-Codex. All models produced substantially higher error rates than the 3.5% reported in the original study; the best result was 10.0% VER with Llama~4~Scout on seaborn. API maturity and stability emerged as key factors in LLM visualization reliability. Seaborn achieved lower error rates, likely because it builds on the well-known matplotlib API, whereas ggplot and Altair showed higher error rates, potentially due to confusion with R's ggplot2 conventions and lower API maturity and adoption. Smaller and cheaper models achieved comparable success rates to larger and more expensive ones, but VER metric alone may be insufficient to determine overall model reliability.
Kokoelmat
  • Opinnäytteet - ylempi korkeakoulututkinto [43236]
Kalevantie 5
PL 617
33014 Tampereen yliopisto
oa[@]tuni.fi | Tietosuoja | Saavutettavuusseloste
 

 

Selaa kokoelmaa

TekijätNimekkeetTiedekunta (2019 -)Tiedekunta (- 2018)Tutkinto-ohjelmat ja opintosuunnatAvainsanatJulkaisuajatKokoelmat

Omat tiedot

Kirjaudu sisäänRekisteröidy
Kalevantie 5
PL 617
33014 Tampereen yliopisto
oa[@]tuni.fi | Tietosuoja | Saavutettavuusseloste