Research Article | Open Access | Download PDF
Volume 74 | Issue 8 | Year 2026 | Article Id. IJETT-V74I8P116 | DOI : https://doi.org/10.14445/22315381/IJETT-V74I8P116Design and Implementation of a Scalable AI-Based Semantic Evaluation System for Hindi Text Using Transformer Models
Nirja D Shah, Jyoti Pareek
| Received | Revised | Accepted | Published |
|---|---|---|---|
| 18 Mar 2026 | 05 Jul 2026 | 22 Jul 2026 | 29 Aug 2026 |
Citation :
Nirja D Shah, Jyoti Pareek, "Design and Implementation of a Scalable AI-Based Semantic Evaluation System for Hindi Text Using Transformer Models," International Journal of Engineering Trends and Technology (IJETT), vol. 74, no. 8, pp. 239-248, 2026. Crossref, https://doi.org/10.14445/22315381/IJETT-V74I8P116
Abstract
Evaluating linguistically diverse descriptive answers in a consistent and accurate manner in modern digital education systems is a growing challenge, especially in low-resource languages like Hindi. Traditional lexical and rule-based grading systems cannot adequately reflect the meaning behind the words, negation, paraphrasing, and so on, which leads to low grading reliability. To overcome these limitations, this study proposes an automated evaluation framework with intelligent rule-based linguistic preprocessing and transformer-based deep learning. The framework uses a fine-tuned multilingual BERT (mBERT) model bhavikardeshna/multilingual-bert-base-cased-hindi to provide contextual embeddings, and cosine similarity-based semantic alignment with the model l3cube-pune/hindi-sentence-similarity-sbert is used to provide automated scores. With optimal setting of learning rate = 5×10⁻⁴, batch size = 24 and epochs = 40, accuracy, precision, recall and F1 score of 78.9%, 80.6%, 77.4% and 79.0% respectively is achieved on HindiRC-Data-master dataset (24 passages, 127 question-answer pairs, grades 2-5) which is more than 14% higher than lexical similarity baselines and is better than previous Hindi QA architectures without domain-specific preprocessing pipelines. The suggested system will save about 40% manual grading, and will enable scalable, consistent and repeatable assessment.
Keywords
Automated answer evaluation, Educational assessment, Hindi Natural Language Processing, Intelligent automation, Semantic similarity, Transformer-based models.
References
[1] Emiliano del
Gobbo et al., “Gradeaid: A Framework for Automatic Short Answers Grading in
Educational Contexts-Design, Implementation and Evaluation,” Knowledge and
Information Systems, vol. 65, no. 10, pp. 4295-4334, 2023.
[CrossRef] [Google Scholar] [Publisher Link]
[2] Khushboo
Khurana et al., “A Textual Question Answering and Handwritten Answer Evaluation
System for the Hindi Language,” International Journal of
Knowledge-based and Intelligent Engineering Systems, vol. 28, no. 3, pp. 166-178, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[3] Shubhankar
Singh et al., “H-AES: Towards Automated Essay Scoring for Hindi,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 13, pp. 15955-15963, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[4] Majdi Beseiso, Omar A. Alzubi, and Hasan
Rashaideh, “A Novel Automated Essay Scoring Approach for Reliable Higher
Educational Assessments,” Journal of Computing in Higher Education, vol.
33, no. 3, pp. 727-746, 2021.
[CrossRef] [Google Scholar] [Publisher Link]
[5] Mustafa Abdul Salam, Mohamed Abd El-Fatah, and
Naglaa Fathy Hassan, “Automatic Grading for Arabic Short Answer Questions using
an Optimized Deep Learning Model,” PLoS ONE, vol. 17, no. 8, pp. 1-41,
2022.
[CrossRef] [Google Scholar] [Publisher Link]
[6] Prerana M.S et al., “Eval -
Automatic Evaluation of Answer Scripts using Deep Learning and Natural Language
Processing,” International Journal of Intelligent Systems and Applications
in Engineering, vol. 11, no. 1, pp. 316-323, 2023.
[Google Scholar] [Publisher Link]
[7] Jacob Devlin et al., “BERT: Pre-Training of Deep
Bidirectional Transformers for Language Understanding,” Proceedings
of the 2019 Conference of the North American Chapter of the Association for
Computational Linguistics: Human Language Technologies, Association for
Computational Linguistics, Minneapolis, Minnesota, vol. 1, pp. 4171-4186, 2019.
[CrossRef] [Google Scholar] [Publisher
Link]
[8] Victor Sanh et al., “DistilBERT, a Distilled
Version of BERT: Smaller, Faster, Cheaper and Lighter,” arXiv, pp. 1-5,
2020.
[CrossRef] [Google Scholar] [Publisher
Link]
[9] Tom Kwiatkowski et al., “Natural Questions: A
Benchmark for Question Answering Research,” Transactions of the Association
for Computational Linguistics, vol. 7, pp. 453-466, 2019.
[CrossRef] [Google Scholar] [Publisher Link]
[10] Anton
Vijeevaraj Ann Sinthusha, Eugene Y.A. Charles, and Ruvan Weerasinghe, “Machine
Reading Comprehension for the Tamil Language with Translated SQuAD,” IEEE
Access, vol. 13, pp. 13312-13328, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[11] Venkatesh Krishnamoorthy
et al., “Evolution of Reading Comprehension and Question Answering Systems,” Procedia
Computer Science, vol. 185, pp. 231-238, 2021.
[CrossRef] [Google Scholar] [Publisher Link]
[12] Deepak Gupta,
Asif Ekbal, and Pushpak Bhattacharyya, “A Deep Neural Network Framework for
English Hindi Question Answering,” ACM Transactions on Asian
and Low-Resource Language Information Processing (TALLIP),
Association for Computing Machinery, New York, United States, vol. 19, no. 2,
pp. 1-22, 2019.
[CrossRef] [Google Scholar] [Publisher
Link]
[13] Somil Gupta,
and Nilesh Khade, “BERT based Multilingual Machine Comprehension in English and
Hindi,” arXiv, pp. 1-13, 2020.
[CrossRef] [Google Scholar] [Publisher
Link]
[14] Pawan Lahoti,
Namita Mittal, and Girdhari Singh, “An Ensemble BERT Model for
English–Hindi–Malayalam Multilingual QA,” Proceedings of Third Emerging
Trends and Technologies on Intelligent Systems, Springer, Singapore, vol.
730, pp. 65-79, 2023.
[CrossRef] [Google Scholar] [Publisher Link]
[15] Nils Reimers, and Iryna
Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” Proceedings
of the 2019 Conference on Empirical Methods in Natural Language Processing and
the 9th International Joint Conference on Natural Language
Processing, Association for Computational Linguistics, Hong Kong, China,
pp. 3982-3992, 2019.
[CrossRef] [Google Scholar] [Publisher
Link]
[16] Simran Khanuja et al.,
“MuRIL: Multilingual Representations for Indian Languages,” arXiv, pp.
1-8, 2021.
[CrossRef] [Google Scholar] [Publisher
Link]
[17] Stefan Haller et al.,
“Survey on Automated Short Answer Grading with Deep Learning: From Word
Embeddings to Transformers,” arXiv, pp. 1-29, 2022.
[CrossRef] [Google Scholar] [Publisher
Link]
[18] Xinhua Zhu,
Han Wu, and Lanfang Zhang, “Automatic Short-Answer Grading via BERT-based Deep Neural
Networks,” IEEE Transactions on Learning Technologies, vol. 15, no. 3,
pp. 364-375, 2022.
[CrossRef] [Google Scholar] [Publisher Link]
[19] Samuel Tobler, “Smart
Grading: A Generative AI-based Tool for Knowledge-Grounded Answer Evaluation in
Educational Assessments,” MethodsX, vol. 12, pp. 1-6, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[20] Li-Hsin Chang, and Filip
Ginter, “Automatic Short Answer Grading for Finnish with ChatGPT,” Proceedings
of the AAAI Conference on Artificial Intelligence, vol. 38, no. 21, pp.
23173-23181, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[21] Ahmad Ayaan, and Kok-Why
Ng, “Automated Grading using Natural Language Processing and Semantic
Analysis,” MethodsX, vol. 14, pp. 1-8, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[22] Sridevi
Bonthu, S. Rama Sree, and M.H.M. Krishna Prasad, “Framework for Automation of Short
Answer Grading based on Domain-Specific Pre-Training,” Engineering
Applications of Artificial Intelligence, vol. 137, 2024.
[CrossRef] [Google Scholar] [Publisher Link]
[23] Raviraj Joshi,
“L3Cube-HindBERT and DevBERT: Pre-Trained BERT Transformer Models for
Devanagari based Hindi and Marathi Languages,” arXiv, pp. 1-7, 2022.
[CrossRef] [Google Scholar] [Publisher
Link]
[24] Kushal Jain et al.,
“Indic-Transformers: An Analysis of Transformer Language Models for Indian
Languages,” arXiv, pp. 1-14, 2020.
[CrossRef] [Google Scholar] [Publisher
Link]
[25] Thomas Wolf et al.,
“Transformers: State-of-the-Art Natural Language Processing,” Proceedings of
the 2020 Conference on Empirical Methods in Natural Language Processing: System
Demonstrations, Association for Computational Linguistics, pp. 38-45, 2020.
[CrossRef] [Google Scholar] [Publisher Link]