HomeUE Research Journalvol. 29 no. 1 (2026)

Automated Essay Evaluator Using Bidirectional Encoder Representations from Transformers Algorithm and Semantic Analysis

Kathleen Dimaano

Discipline: others in technology

 

Abstract:

This study introduces the design and assessment of the Automated Essay Evaluator (AEE) system based on the use of the Bidirectional Encoder Representations from Transformers (BERT) algorithm and Semantic Analysis, which will help overcome some problems encountered with conventional essay evaluation methods. Using a developmental descriptive research design guided by the CRISP-DM process, the system integrates multiple Natural Language Processing models to assess grammar, structure, semantic relevance, and originality with a high degree of accuracy. Designed for use in educational institutions, the AEE integrates advanced Natural Language Processing (NLP) models such as BERT-base-uncasedd, all-MiniLM-L6-v2, spelling-correctionEnglish-based, chatgpt-detector-roberta, and ms-macro-MiniLM-L6-v2 to assess grammar, structure, semantic relevance, and originality. The AEE system ensures wide-ranging evaluation because it determines if there are semantic relationships between essays, identifies grammatical errors, checks for plagiarism, and detects if the content is artificial intelligence (AI) generated. Statistical analysis revealed a strong linear association between the human consensus and automated scoring through Pearson’s r = 0.9700 where p < 0.001. To verify fairness, a two-way ANOVA confirmed no statistically significant difference between human and automated scoring methods with value of p = 0.297, while the Intraclass Correlation Coefficient (ICC 2,1) yielded a value of 0.962, indicating excellent reliability. These results demonstrate that the application’s Automated Scoring driven by semantic validation. The result demonstrates high accuracy when it comes to plagiarism and AI generated detection with a score of 94%. The ISO/IEC 25010 criteria for software quality factors were used, and they received very good ratings, especially in the Safety category with a mean weighted score of 4.86. The flexibility, adaptability, and usability aspects of the system were obtained through feedback from the experts who are teachers, students, and information technology practitioners.



References:

  1. Agarwal, S., & Meena, S. (2021). Paraphrased plagiarism detection using semantic similarity based on transformer networks. Journal of Ambient Intelligence and Humanized Computing, 12(11), 9987-10000.
  2. Bates, T., Cobo, C., Mariño, O., & Wheeler, S. (2020). Can artificial intelligence transform higher education? International Journal of Educational Technology in Higher Education, 17(1). https://doi.org/10.1186/s41239-020-00218-x
  3. Bellini, V., Semeraro, F., Montomoli, J., Cascella, M., & Bignami, E. (2024). Between human and AI: assessing the reliability of AI text detection tools. Current Medical Research and Opinion, 40(3), 353–358.
  4. Chatti, M. A., Muslim, A., Guliani, M., & Guesmi, M. (2020). The LAVA model: Learning analytics meets visual analytics. In D. Ifenthaler & D. C. Gibson (Eds.), Adoption of data analytics in higher education learning and teaching (pp. 70–93). Springer. https://doi.org/10.1007/978-3-030-47392-1_5
  5. Chen, T., & Du, L. (2022). English Semantic Analysis Algorithm and application based on improved attention Mechanism model. Mathematical Problems in Engineering, 2022, 1–9. https://doi.org/10.1155/2022/2165537
  6. Chumbar, S. (2023, September 24). The CRISP-DM Process: A Comprehensive Guide. Medium. https://medium.com/@shawn.chumbar/the-crisp-dm-process-a-comprehensive-guide-4d893aecb151
  7. Cox, A. M. (2021). Exploring the impact of Artificial Intelligence and robots on higher education through literature-based design fictions. International Journal of Educational Technology in Higher Education, 18(1). https://doi.org/10.1186/s41239-020-00237-8
  8. De Laat, M., Joksimovic, S., & Ifenthaler, D. (2020). Artificial intelligence, real-time feedback and workplace learning analytics to support in situ complex problem-solving: A commentary. The International Journal of Information and Learning Technology, 37(5), 267–277. https://doi.org/10.1108/IJILT-03-2020-0026
  9. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT (pp. 4171–4186). Association for Computational Linguistics.
  10. Du, C., Zhang, Y., & Liu, L. (2023). Sentence-BERT Based Off-Topic Composition Detection Algorithm. Proceedings of the 2023 4th International Conference on Computing, Networks and Internet of Things. https://doi.org/10.1145/3603781.3603871
  11. EDUCAUSE Review. (2023, November). Academic Integrity in the Age of AI. Retrieved from https://er.educause.edu/articles/sponsored/2023/11/ academic-integrity-in-the-age-of-ai
  12. Erol, G., Ergen, A., Gülşen Erol, B., Kaya Ergen, Ş., Bora, T. S., Çölgeçen, A. D., ... & Güngör, A. (2025). Can we trust academic AI detectors? Accuracy and limitations of AI-output detectors. Acta Neurochirurgica, 167(1), 214.
  13. Fang, M., Hutson, A. D., & Yu, H. (2025). Robust permutation test of intraclass correlation coefficient for assessing agreement. Cancers, 17(16), 2713.
  14. Gabdullina, S. (2023). THE POTENTIAL OF THE ESSAY IN FORMATIVE ASSESSMENT: LITERATURE REVIEW. Education. Innovation. Diversity, 1(6), 48–53. https://doi.org/10.17770/eid2023.1.7176
  15. Garrido-Merchan, E. C., Gozalo-Brizuela, R., & Gonzalez-Carvajal, S. (2023). Comparing BERT against Traditional Machine Learning Models in Text Classification. Journal of Computational and Cognitive Engineering. https://doi.org/10.47852/bonviewjcce3202838
  16. Ghosal, S. S., Chakraborty, S., Geiping, J., Huang, F., Manocha, D., & Bedi, A. (2023, December 30). A survey on the Possibilities & Impossibilities of AI-generated text Detection. OpenReview. https://openreview.net/ forum?id=AXtFeYjboj
  17. Gruetzemacher, R. (2022, April 19). The Power of Natural Language Processing. Harvard Business Review. https://hbr.org/2022/04/the-power-of-natural-language-processing
  18. Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22(1), 4.
  19. Hussein, R. R., Razak, Z. A., & Hashim, H. S. (2023). Automated and handmade features in automated essay evaluation (AEE): A systematic literature review. Education and Information Technologies, 28(1), 1–36.
  20. Ifenthaler, D. (2022). Automated Essay Scoring Systems. Handbook of Open, Distance and Digital Education, 1–15. https://doi.org/10.1007/978-981-19-0351-9_59-1
  21. Ifenthaler, D., & Schumacher, C. (2023). Reciprocal issues of artificial and human intelligence in education. Journal of Research on Technology in Education, 55(1), 1–6. https://doi.org/10.1080/15391523.2022.2154511
  22. Ippolito, D., Fischer, M., Wang, Z., & De Cao, N. (2020). Automatic detection of machine-generated text: Challenges and opportunities. arXiv preprint arXiv:2003.06851.
  23. Jiffriya, M., Jahan, M., & Ragel, R. (2021). Plagiarism Detection Tools and Techniques: A Comprehensive survey. Journal of Science-FAS-SEUSL (2021), 02(02) 47-64(2738–2184), 11358. https://seu.ac.lk/jsc/ publication/v2n2/Manuscript%205.pdf
  24. Kanade, V. (2022, June 16). What Is Semantic Analysis? Definition, Examples, and Applications in 2022 |. Spiceworks. https://www.spiceworks.com/tech/artificial-intelligence/articles/what-is-semantic-analysis/
  25. Karoo, K., & Meghraj Jogi, M. (2023). International Journal of Research Publication and Reviews Syntactic and Semantic Analysis in Natural Language Processing: Unveiling the Underlying Mechanisms. International Journal of Research Publication and Reviews, 4, 381–395. https://ijrpr.com/uploads/V4ISSUE12/IJRPR20061.pdf
  26. Kaur, S., & Gupta, P. (2020). An Empirical Analysis of BERT Embedding for Automated Essay Scoring. ResearchGate. https://www.researchgate.net/ publication/346085252
  27. Khan, W., Daud, A., Khan, K., Muhammad, S., & Haq, R. (2023). Exploring the frontiers of deep learning and natural language processing: A comprehensive overview of key challenges and emerging trends. Natural Language Processing Journal, 4, 100026. https://doi.org/ 10.1016/j.nlp.2023.100026
  28. Kim, J. K., Chua, M., Rickard, M., & Lorenzo, A. (2023). ChatGPT and large language model (LLM) chatbots: the current state of acceptability and a proposal for guidelines on utilization in academic medicine. Journal of Pediatric Urology.
  29. Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. https://doi.org/10.1016/j.jcm.2016.02.012
  30. Kumar, R., & Boulanger, D. (2020). Enhancing automated essay scoring using grammar and syntax-aware features. International Journal of Artificial Intelligence in Education, 30(2), 243–260. https://doi.org/10.1007/ s40593-020-00196-y
  31. Kumar, V., & Boulanger, D. (2020). Explainable Automated Essay Scoring: Deep Learning Really Has Pedagogical Value. Frontiers in Education, 5. https://doi.org/10.3389/feduc.2020.572367
  32. Kusuma, J. S., Halim, K., Pranoto, E. J. P., & Kanigoro, B. (2022). Automated essay scoring using machine learning. Proceedings of the 2022 International Conference on Cybernetics and Intelligent Systems. https://doi.org/10.1109/ICORIS56080.2022.10031338
  33. Lancaster, T., & Cullinan, J. (2019). AI and the future of assessment: Academic integrity beyond the essay. Journal of Academic Ethics, 17(2), 113-128.
  34. Lewis Sevcikova, B. (2018). Human versus Automated Essay Scoring: A Critical Review. Arab World English Journal, 9(2), 157–174. https://doi.org/10.24093/awej/vol9no2.11
  35. Li, S. (2024). Automated essay scoring: Recent successes and future directions. Automated Essay Scoring: A Reflection on the State of the Art (survey/position paper). ACL/EMNLP 2024 proceedings. https://aclanthology.org/2024.emnlp-main.991.pdf
  36. Liu, S., Zhang, Y., & Cao, J. (2024). BERT-Enhanced Retrieval and Plagiarism Identification Based on Faiss. arXiv Preprint. https://arxiv.org/ abs/2404.01582
  37. Mah, P. M., Skalna, I., & Muzam, J. (2022). Natural Language Processing and Artificial Intelligence for Enterprise Management in the Era of Industry 4.0. Applied Sciences, 12(18), 9207. https://doi.org/10.3390/app12189207
  38. Misgna, H. (2024). A survey on deep learning–based automated essay scoring models. International Journal of Intelligent Systems / Springer. https://doi.org/10.1007/s10462-024-11017-5
  39. Narkhede, S. (2021, June 15). Understanding Confusion Matrix - towards Data science. Medium. https://towardsdatascience.com/understanding confusion-matrix-a9ad42dcfd62
  40. Nur, M., Ramadhan, A., & Hendric, L. (2023). Automatic essay exam scoring system: a systematic literature review. Procedia Computer Science, 216, 531–538. https://doi.org/10.1016/j.procs.2022.12.166
  41. Oben, A. I. (2021). Research instruments: A Questionnaire and an interview guide used to investigate  the implementation of higher education objectives and the attainment of Cameroon’s vision 2035. European Journal of Education Studies, 8(7). https://doi.org/10.46827/ejes.v8i7.3808
  42. Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education. Artificial Intelligence, 6, 100234–100234. https://doi.org/10.1016/j.caeai.2024.100234
  43. Palermo, C., et al. (2022). Rater characteristics, response content, and scoring contexts: An integrated examination of rater accuracy. Frontiers in Psychology, 13. https://doi.org/10.3389/fpsyg.2022.937097
  44. Premalatha, M., Viswanathan, V., & Čepová, L. (2022). Application of semantic Analysis and LSTM-GRU in developing a personalized course recommendation system. Applied Sciences, 12(21), 10792. https://doi.org/10.3390/app122110792
  45. Putri, N. M., Hidayatullah, S., & Nugroho, R. P. (2023). The Impact of Sentence Tokenization in Indonesian Automated Essay Scoring. Research in Adaptive and Innovative Algorithms, 37(5), 562–570. https://doi.org/10.18280/ria.370502
  46. Quidwai, M. A., Li, C., & Dube, P. (2023). Beyond Black Box AI-Generated Plagiarism Detection: From Sentence to Document Level. arXiv preprint arXiv:2306.08122. https://arxiv.org/abs/2306.08122
  47. Ramesh, D., & Sanampudi, S. K. (2021). An automated essay scoring systems: a systematic literature review. Artificial Intelligence Review, 55, 2495–2527. https://doi.org/10.1007/s10462-021-10068-2
  48. Ramesh, D., Sanampudi, B., & others. (2022). An automated essay scoring systems: A systematic literature review. Applied Intelligence (Springer). https://doi.org/10.1007/s10462-021-10068-2
  49. Robeson, S. M., & Willmott, C. J. (2023). Decomposition of the mean absolute error (MAE) into systematic and unsystematic components. PLOS ONE, 18(2), e0279774. https://doi.org/10.1371/journal.pone.0279774
  50. Scribbr. (2024). The Best Grammar Checker? Here’s Our Top 10.
  51. Shi Huawei, & Aryadoust, V. (2025). Accuracy of current automated essay evaluation (AEE) systems in English writing assessment: A systematic review. Computers and Education Open, 7, 100203.
  52. Sternberg, R. J. (2023, November 18). human intelligence. Encyclopedia Britannica.
  53. Taber, K. S. (2017). The use of Cronbach’s Alpha when developing and reporting research instruments in science education. Research in Science Education, 48(6), 1273–1296. https://doi.org/10.1007/s11165-016-9602-2
  54. Ten Hove, D., Jorgensen, T. D., & van der Ark, L. (2024). Updated guidelines on selecting an intraclass correlation coefficient for interrater reliability, with applications to incomplete observational designs. Psychological Methods. https://doi.org/10.1037/met0000516
  55. Uto, M. (2021). A review of deep-neural automated essay scoring models. Behaviormetrika, 48(2), 459–484. https://doi.org/10.1007/s41237-021-00142-y
  56. Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., & Zhou, M. (2020). MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. arXiv preprint arXiv:2002.10957. https://doi.org/10.48550/ arXiv.2002.10957
  57. Wang, Y., & Litman, D. (2022). Evaluation of Semantic Similarity Tools for Automated Content Scoring of EFL Essays. Education and Information Technologies. https://doi.org/10.1007/s10639-022-11179-1
  58. Wang, Y., Wang, C., Li, R., & Lin, H. (2022). On the use of BERT for automated essay scoring: Joint learning of multi-scale essay representation. In Proceedings of NAACL 2022. https://aclanthology.org/2022.naacl-main.249/
  59. Xian, J., Yuan, J., Zheng, P., Chen, D., & Yuntao, N. (2024). BERT-Enhanced Retrieval Tool for Homework Plagiarism Detection System. arXiv preprint arXiv:2404.01582. https://arxiv.org/abs/2404.01582
  60. Yousif, M. J. (2023). Systematic Review of Semantic Analysis Methods. Applied Computing Journal, 286–300. https://doi.org/10.52098/acj.2023346
  61. Zhao, W. (2022). Inspired, but not mimicking: a conversation between artificial intelligence and human intelligence. National Science Review, 9(6). https://doi.org/10.1093/nsr/nwac068
  62. Zhao, X., Xu, J., & Liu, X. (2022). Multi-Scale Essay Representation with BERT in AES. Proceedings of the NAACL 2022 Conference. https://aclanthology.org/2022.naacl-main.249