GRIS - Grammatical Interpretable Translation Scoring : An Interpretable Dependency-Based Metric for Machine Translation Evaluation
Lumbini, Palada Pathirage Chalitha (2026)
Lumbini, Palada Pathirage Chalitha
2026
Matematiikan ja tilastollisen data-analyysin maisteriohjelma - Master's Programme in Mathematics and Statistical Data Analytics
Informaatioteknologian ja viestinnän tiedekunta - Faculty of Information Technology and Communication Sciences
This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.
Hyväksymispäivämäärä
2026-06-15
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202606127347
https://urn.fi/URN:NBN:fi:tuni-202606127347
Tiivistelmä
The automatic evaluation of machine translation (MT) quality remains a fundamental challenge in natural language processing. Traditional metrics such as BLEU provide fast and reproducible results based on surface level n-gram overlap, but often fail to capture similarities in meaning and sentence structure. Recent neural metrics such as BERTScore and COMET achieve higher correlation with human judgements, but operate as black-box models with limited interpretability, making it difficult to diagnose specific translation errors.
This thesis introduces GRIS (Grammatical Interpretable Translation Scoring), a transparent framework for machine translation evaluation designed to improve interpretability and reliability of automatic translation evaluation. This approach evaluates the translations using two components: structural similarity and semantic similarity. Structural similarity is computed using Universal Dependencies to compare grammatical relations between hypothesis and reference sentences. Semantic similarity measured using contextual sentence embeddings. These components are combined into a single score that captures both semantic accuracy and grammatical correctness.
GRIS consists of two complementary methods. GRIS-DepScore evaluates translations using dependency edge relations, while GRIS-SynGram evaluates translation using subtree path syntactic n-gram matching. Both methods rely on Universal Dependencies to model grammatical structure and apply linguistic rules to capture features such as passive voice and negation.
The effectiveness of GRIS was evaluated on multilingual machine translation datasets using WMT22 MQM data for Chinese–English (zh→en), English–German (en→de), and English–Russian (en→ru) translation tasks. Experimental results show that GRIS-DepScore outperforms BERTScore at system level for en→de and en→ru, and at segment level for en→ru. GRIS-SynGram also outperforms BERTScore at system level for en→ru. This demonstrates the effectiveness of structure based evaluation for morphologically rich languages without requiring source sentences or training data.
The key contribution of GRIS is its explainable evaluation analysis of translation quality and linguistic errors. Sentence level penalties are applied for major structural mismatches, while local penalties capture errors such as incorrect argument roles and missing modal verbs. This design makes GRIS useful for machine translation error analysis and system diagnosis. Second, it analyses how language structure affects evaluation performance. It shows that the usefulness of structure based evaluation methods changes depending on the grammatical characteristics of the languages being translated.
This thesis introduces GRIS (Grammatical Interpretable Translation Scoring), a transparent framework for machine translation evaluation designed to improve interpretability and reliability of automatic translation evaluation. This approach evaluates the translations using two components: structural similarity and semantic similarity. Structural similarity is computed using Universal Dependencies to compare grammatical relations between hypothesis and reference sentences. Semantic similarity measured using contextual sentence embeddings. These components are combined into a single score that captures both semantic accuracy and grammatical correctness.
GRIS consists of two complementary methods. GRIS-DepScore evaluates translations using dependency edge relations, while GRIS-SynGram evaluates translation using subtree path syntactic n-gram matching. Both methods rely on Universal Dependencies to model grammatical structure and apply linguistic rules to capture features such as passive voice and negation.
The effectiveness of GRIS was evaluated on multilingual machine translation datasets using WMT22 MQM data for Chinese–English (zh→en), English–German (en→de), and English–Russian (en→ru) translation tasks. Experimental results show that GRIS-DepScore outperforms BERTScore at system level for en→de and en→ru, and at segment level for en→ru. GRIS-SynGram also outperforms BERTScore at system level for en→ru. This demonstrates the effectiveness of structure based evaluation for morphologically rich languages without requiring source sentences or training data.
The key contribution of GRIS is its explainable evaluation analysis of translation quality and linguistic errors. Sentence level penalties are applied for major structural mismatches, while local penalties capture errors such as incorrect argument roles and missing modal verbs. This design makes GRIS useful for machine translation error analysis and system diagnosis. Second, it analyses how language structure affects evaluation performance. It shows that the usefulness of structure based evaluation methods changes depending on the grammatical characteristics of the languages being translated.
