Toward AI Evaluation of Student Essays
Rantanen, Petri; Saari, Mika; Virta, Ulla-Talvikki; Abrahamsson, Pekka (2025)
Rantanen, Petri
Saari, Mika
Virta, Ulla-Talvikki
Abrahamsson, Pekka
2025
This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202601151471
https://urn.fi/URN:NBN:fi:tuni-202601151471
Kuvaus
Peer reviewed
Tiivistelmä
The integration of artificial intelligence (AI) into teaching practices has expanded rapidly, and there is increasing discussion about computer-assisted assessment. This study investigates the application of AI for evaluating student essays, focusing on a Software Testing course at Tampere University. Traditionally, on this course essays are assessed manually by teaching assistants using predefined evaluation criteria. This research explores the feasibility of automating the evaluation process by comparing assessments generated by AI models - OpenAI's large language models (LLM) API - with human evaluations. The study focuses on the research question: Can automated assessment replace human evaluation? The key stages in this study include selecting a suitable AI model, its implementation and configuration, and designing an automated assessment workflow that aligns with human evaluation. The results indicate that, while AI models can evaluate at a level comparable to humans, current language models have limitations, such as consistency issues and an inability to interpret context. The findings of this study suggest that AI can be used as complementary tools rather than fully replacing humans. This work highlights the potential of AI in educational assessment while acknowledging the need for further refinement in model performance. Future directions include exploring open-source models and addressing the technical challenges related to resource requirements.
Kokoelmat
- TUNICRIS-julkaisut [24991]
