Automated Essay Evaluator Using Bidirectional Encoder Representations from Transformers Algorithm and Semantic Analysis
Kathleen Dimaano
Discipline: others in technology
Abstract:
This study introduces the design and assessment of the Automated Essay
Evaluator (AEE) system based on the use of the Bidirectional Encoder
Representations from Transformers (BERT) algorithm and Semantic Analysis,
which will help overcome some problems encountered with conventional essay
evaluation methods. Using a developmental descriptive research design guided by
the CRISP-DM process, the system integrates multiple Natural Language
Processing models to assess grammar, structure, semantic relevance, and
originality with a high degree of accuracy. Designed for use in educational
institutions, the AEE integrates advanced Natural Language Processing (NLP)
models such as BERT-base-uncasedd, all-MiniLM-L6-v2, spelling-correctionEnglish-based, chatgpt-detector-roberta, and ms-macro-MiniLM-L6-v2 to assess
grammar, structure, semantic relevance, and originality. The AEE system ensures
wide-ranging evaluation because it determines if there are semantic relationships
between essays, identifies grammatical errors, checks for plagiarism, and detects
if the content is artificial intelligence (AI) generated. Statistical analysis revealed a
strong linear association between the human consensus and automated scoring
through Pearson’s r = 0.9700 where p < 0.001. To verify fairness, a two-way
ANOVA confirmed no statistically significant difference between human and
automated scoring methods with value of p = 0.297, while the Intraclass
Correlation Coefficient (ICC 2,1) yielded a value of 0.962, indicating excellent
reliability. These results demonstrate that the application’s Automated Scoring
driven by semantic validation. The result demonstrates high accuracy when it
comes to plagiarism and AI generated detection with a score of 94%. The ISO/IEC
25010 criteria for software quality factors were used, and they received very good
ratings, especially in the Safety category with a mean weighted score of 4.86. The
flexibility, adaptability, and usability aspects of the system were obtained through
feedback from the experts who are teachers, students, and information technology
practitioners.
References:
- Agarwal, S., & Meena, S. (2021). Paraphrased plagiarism detection using semantic similarity based on transformer networks. Journal of Ambient Intelligence and Humanized Computing, 12(11), 9987-10000.
- Bates, T., Cobo, C., Mariño, O., & Wheeler, S. (2020). Can artificial intelligence transform higher education? International Journal of Educational Technology in Higher Education, 17(1). https://doi.org/10.1186/s41239-020-00218-x
- Bellini, V., Semeraro, F., Montomoli, J., Cascella, M., & Bignami, E. (2024). Between human and AI: assessing the reliability of AI text detection tools. Current Medical Research and Opinion, 40(3), 353–358.
- Chatti, M. A., Muslim, A., Guliani, M., & Guesmi, M. (2020). The LAVA model: Learning analytics meets visual analytics. In D. Ifenthaler & D. C. Gibson (Eds.), Adoption of data analytics in higher education learning and teaching (pp. 70–93). Springer. https://doi.org/10.1007/978-3-030-47392-1_5
- Chen, T., & Du, L. (2022). English Semantic Analysis Algorithm and application based on improved attention Mechanism model. Mathematical Problems in Engineering, 2022, 1–9. https://doi.org/10.1155/2022/2165537
- Chumbar, S. (2023, September 24). The CRISP-DM Process: A Comprehensive Guide. Medium. https://medium.com/@shawn.chumbar/the-crisp-dm-process-a-comprehensive-guide-4d893aecb151
- Cox, A. M. (2021). Exploring the impact of Artificial Intelligence and robots on higher education through literature-based design fictions. International Journal of Educational Technology in Higher Education, 18(1). https://doi.org/10.1186/s41239-020-00237-8
- De Laat, M., Joksimovic, S., & Ifenthaler, D. (2020). Artificial intelligence, real-time feedback and workplace learning analytics to support in situ complex problem-solving: A commentary. The International Journal of Information and Learning Technology, 37(5), 267–277. https://doi.org/10.1108/IJILT-03-2020-0026
- Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT (pp. 4171–4186). Association for Computational Linguistics.
- Du, C., Zhang, Y., & Liu, L. (2023). Sentence-BERT Based Off-Topic Composition Detection Algorithm. Proceedings of the 2023 4th International Conference on Computing, Networks and Internet of Things. https://doi.org/10.1145/3603781.3603871
- EDUCAUSE Review. (2023, November). Academic Integrity in the Age of AI. Retrieved from https://er.educause.edu/articles/sponsored/2023/11/ academic-integrity-in-the-age-of-ai
- Erol, G., Ergen, A., Gülşen Erol, B., Kaya Ergen, Ş., Bora, T. S., Çölgeçen, A. D., ... & Güngör, A. (2025). Can we trust academic AI detectors? Accuracy and limitations of AI-output detectors. Acta Neurochirurgica, 167(1), 214.
- Fang, M., Hutson, A. D., & Yu, H. (2025). Robust permutation test of intraclass correlation coefficient for assessing agreement. Cancers, 17(16), 2713.
- Gabdullina, S. (2023). THE POTENTIAL OF THE ESSAY IN FORMATIVE ASSESSMENT: LITERATURE REVIEW. Education. Innovation. Diversity, 1(6), 48–53. https://doi.org/10.17770/eid2023.1.7176
- Garrido-Merchan, E. C., Gozalo-Brizuela, R., & Gonzalez-Carvajal, S. (2023). Comparing BERT against Traditional Machine Learning Models in Text Classification. Journal of Computational and Cognitive Engineering. https://doi.org/10.47852/bonviewjcce3202838
- Ghosal, S. S., Chakraborty, S., Geiping, J., Huang, F., Manocha, D., & Bedi, A. (2023, December 30). A survey on the Possibilities & Impossibilities of AI-generated text Detection. OpenReview. https://openreview.net/ forum?id=AXtFeYjboj
- Gruetzemacher, R. (2022, April 19). The Power of Natural Language Processing. Harvard Business Review. https://hbr.org/2022/04/the-power-of-natural-language-processing
- Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22(1), 4.
- Hussein, R. R., Razak, Z. A., & Hashim, H. S. (2023). Automated and handmade features in automated essay evaluation (AEE): A systematic literature review. Education and Information Technologies, 28(1), 1–36.
- Ifenthaler, D. (2022). Automated Essay Scoring Systems. Handbook of Open, Distance and Digital Education, 1–15. https://doi.org/10.1007/978-981-19-0351-9_59-1
- Ifenthaler, D., & Schumacher, C. (2023). Reciprocal issues of artificial and human intelligence in education. Journal of Research on Technology in Education, 55(1), 1–6. https://doi.org/10.1080/15391523.2022.2154511
- Ippolito, D., Fischer, M., Wang, Z., & De Cao, N. (2020). Automatic detection of machine-generated text: Challenges and opportunities. arXiv preprint arXiv:2003.06851.
- Jiffriya, M., Jahan, M., & Ragel, R. (2021). Plagiarism Detection Tools and Techniques: A Comprehensive survey. Journal of Science-FAS-SEUSL (2021), 02(02) 47-64(2738–2184), 11358. https://seu.ac.lk/jsc/ publication/v2n2/Manuscript%205.pdf
- Kanade, V. (2022, June 16). What Is Semantic Analysis? Definition, Examples, and Applications in 2022 |. Spiceworks. https://www.spiceworks.com/tech/artificial-intelligence/articles/what-is-semantic-analysis/
- Karoo, K., & Meghraj Jogi, M. (2023). International Journal of Research Publication and Reviews Syntactic and Semantic Analysis in Natural Language Processing: Unveiling the Underlying Mechanisms. International Journal of Research Publication and Reviews, 4, 381–395. https://ijrpr.com/uploads/V4ISSUE12/IJRPR20061.pdf
- Kaur, S., & Gupta, P. (2020). An Empirical Analysis of BERT Embedding for Automated Essay Scoring. ResearchGate. https://www.researchgate.net/ publication/346085252
- Khan, W., Daud, A., Khan, K., Muhammad, S., & Haq, R. (2023). Exploring the frontiers of deep learning and natural language processing: A comprehensive overview of key challenges and emerging trends. Natural Language Processing Journal, 4, 100026. https://doi.org/ 10.1016/j.nlp.2023.100026
- Kim, J. K., Chua, M., Rickard, M., & Lorenzo, A. (2023). ChatGPT and large language model (LLM) chatbots: the current state of acceptability and a proposal for guidelines on utilization in academic medicine. Journal of Pediatric Urology.
- Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. https://doi.org/10.1016/j.jcm.2016.02.012
- Kumar, R., & Boulanger, D. (2020). Enhancing automated essay scoring using grammar and syntax-aware features. International Journal of Artificial Intelligence in Education, 30(2), 243–260. https://doi.org/10.1007/ s40593-020-00196-y
- Kumar, V., & Boulanger, D. (2020). Explainable Automated Essay Scoring: Deep Learning Really Has Pedagogical Value. Frontiers in Education, 5. https://doi.org/10.3389/feduc.2020.572367
- Kusuma, J. S., Halim, K., Pranoto, E. J. P., & Kanigoro, B. (2022). Automated essay scoring using machine learning. Proceedings of the 2022 International Conference on Cybernetics and Intelligent Systems. https://doi.org/10.1109/ICORIS56080.2022.10031338
- Lancaster, T., & Cullinan, J. (2019). AI and the future of assessment: Academic integrity beyond the essay. Journal of Academic Ethics, 17(2), 113-128.
- Lewis Sevcikova, B. (2018). Human versus Automated Essay Scoring: A Critical Review. Arab World English Journal, 9(2), 157–174. https://doi.org/10.24093/awej/vol9no2.11
- Li, S. (2024). Automated essay scoring: Recent successes and future directions. Automated Essay Scoring: A Reflection on the State of the Art (survey/position paper). ACL/EMNLP 2024 proceedings. https://aclanthology.org/2024.emnlp-main.991.pdf
- Liu, S., Zhang, Y., & Cao, J. (2024). BERT-Enhanced Retrieval and Plagiarism Identification Based on Faiss. arXiv Preprint. https://arxiv.org/ abs/2404.01582
- Mah, P. M., Skalna, I., & Muzam, J. (2022). Natural Language Processing and Artificial Intelligence for Enterprise Management in the Era of Industry 4.0. Applied Sciences, 12(18), 9207. https://doi.org/10.3390/app12189207
- Misgna, H. (2024). A survey on deep learning–based automated essay scoring models. International Journal of Intelligent Systems / Springer. https://doi.org/10.1007/s10462-024-11017-5
- Narkhede, S. (2021, June 15). Understanding Confusion Matrix - towards Data science. Medium. https://towardsdatascience.com/understanding confusion-matrix-a9ad42dcfd62
- Nur, M., Ramadhan, A., & Hendric, L. (2023). Automatic essay exam scoring system: a systematic literature review. Procedia Computer Science, 216, 531–538. https://doi.org/10.1016/j.procs.2022.12.166
- Oben, A. I. (2021). Research instruments: A Questionnaire and an interview guide used to investigate the implementation of higher education objectives and the attainment of Cameroon’s vision 2035. European Journal of Education Studies, 8(7). https://doi.org/10.46827/ejes.v8i7.3808
- Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education. Artificial Intelligence, 6, 100234–100234. https://doi.org/10.1016/j.caeai.2024.100234
- Palermo, C., et al. (2022). Rater characteristics, response content, and scoring contexts: An integrated examination of rater accuracy. Frontiers in Psychology, 13. https://doi.org/10.3389/fpsyg.2022.937097
- Premalatha, M., Viswanathan, V., & Čepová, L. (2022). Application of semantic Analysis and LSTM-GRU in developing a personalized course recommendation system. Applied Sciences, 12(21), 10792. https://doi.org/10.3390/app122110792
- Putri, N. M., Hidayatullah, S., & Nugroho, R. P. (2023). The Impact of Sentence Tokenization in Indonesian Automated Essay Scoring. Research in Adaptive and Innovative Algorithms, 37(5), 562–570. https://doi.org/10.18280/ria.370502
- Quidwai, M. A., Li, C., & Dube, P. (2023). Beyond Black Box AI-Generated Plagiarism Detection: From Sentence to Document Level. arXiv preprint arXiv:2306.08122. https://arxiv.org/abs/2306.08122
- Ramesh, D., & Sanampudi, S. K. (2021). An automated essay scoring systems: a systematic literature review. Artificial Intelligence Review, 55, 2495–2527. https://doi.org/10.1007/s10462-021-10068-2
- Ramesh, D., Sanampudi, B., & others. (2022). An automated essay scoring systems: A systematic literature review. Applied Intelligence (Springer). https://doi.org/10.1007/s10462-021-10068-2
- Robeson, S. M., & Willmott, C. J. (2023). Decomposition of the mean absolute error (MAE) into systematic and unsystematic components. PLOS ONE, 18(2), e0279774. https://doi.org/10.1371/journal.pone.0279774
- Scribbr. (2024). The Best Grammar Checker? Here’s Our Top 10.
- Shi Huawei, & Aryadoust, V. (2025). Accuracy of current automated essay evaluation (AEE) systems in English writing assessment: A systematic review. Computers and Education Open, 7, 100203.
- Sternberg, R. J. (2023, November 18). human intelligence. Encyclopedia Britannica.
- Taber, K. S. (2017). The use of Cronbach’s Alpha when developing and reporting research instruments in science education. Research in Science Education, 48(6), 1273–1296. https://doi.org/10.1007/s11165-016-9602-2
- Ten Hove, D., Jorgensen, T. D., & van der Ark, L. (2024). Updated guidelines on selecting an intraclass correlation coefficient for interrater reliability, with applications to incomplete observational designs. Psychological Methods. https://doi.org/10.1037/met0000516
- Uto, M. (2021). A review of deep-neural automated essay scoring models. Behaviormetrika, 48(2), 459–484. https://doi.org/10.1007/s41237-021-00142-y
- Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., & Zhou, M. (2020). MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. arXiv preprint arXiv:2002.10957. https://doi.org/10.48550/ arXiv.2002.10957
- Wang, Y., & Litman, D. (2022). Evaluation of Semantic Similarity Tools for Automated Content Scoring of EFL Essays. Education and Information Technologies. https://doi.org/10.1007/s10639-022-11179-1
- Wang, Y., Wang, C., Li, R., & Lin, H. (2022). On the use of BERT for automated essay scoring: Joint learning of multi-scale essay representation. In Proceedings of NAACL 2022. https://aclanthology.org/2022.naacl-main.249/
- Xian, J., Yuan, J., Zheng, P., Chen, D., & Yuntao, N. (2024). BERT-Enhanced Retrieval Tool for Homework Plagiarism Detection System. arXiv preprint arXiv:2404.01582. https://arxiv.org/abs/2404.01582
- Yousif, M. J. (2023). Systematic Review of Semantic Analysis Methods. Applied Computing Journal, 286–300. https://doi.org/10.52098/acj.2023346
- Zhao, W. (2022). Inspired, but not mimicking: a conversation between artificial intelligence and human intelligence. National Science Review, 9(6). https://doi.org/10.1093/nsr/nwac068
- Zhao, X., Xu, J., & Liu, X. (2022). Multi-Scale Essay Representation with BERT in AES. Proceedings of the NAACL 2022 Conference. https://aclanthology.org/2022.naacl-main.249
Full Text:
Note: Kindly Login or Register to gain access to this article.
ISSN 3116-5354 (Online)
ISSN 3028-0435 (Print)