Guidelines For Automatic Grading of Student Essays Using Large Language Models
DOI:
https://doi.org/10.55549/epess.1037Keywords:
Higher education, Critical-thinking, Automated grading, Large language models, Task distribution, Ensemble LLMAbstract
Automated essay evaluation using large language models (LLMs) has emerged as a promising approach to support scalable and consistent educational assessment. However, the effectiveness of LLM-based grading varies significantly across evaluation dimensions and is highly influenced by prompt design and model selection. In this study, we evaluate five state-of-the-art LLMs across five rubric-based categories: Relevance to Question, Reasoning and Critical Thinking, Evidence and Examples, Organization, and Clarity and Writing Quality. We systematically investigate the impact of three prompting strategies, including rubric-only prompting, exemplar-based prompting (with and without rubric guidance)(Original and Refined prompt designs) incorporating structured instructions. Additionally, a prompt ablation study is conducted to analyze the effect of instruction detail and reasoning guidance on grading accuracy. Our results demonstrate that no single model consistently outperforms others across all categories. Instead, each LLM exhibits strengths in specific evaluation dimensions, motivating a category-wise model selection approach. Furthermore, refined and structured prompting strategies significantly improve evaluation performance, with chain-of-thought-style instructions yielding the highest accuracy. These findings highlight the importance of both prompt engineering and task-specialized model allocation in developing robust LLM-based grading systems. The study provides evidence supporting multi-agent frameworks, where different LLMs can be assigned to distinct evaluation tasks to enhance overall grading reliability and alignment with human judgments.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 The Eurasia Proceedings of Educational and Social Sciences

This work is licensed under a Creative Commons Attribution 4.0 International License.
The articles may be used for research, teaching, and private study purposes. Any substantial or systematic reproduction, redistribution, reselling, loan, sub-licensing, systematic supply, or distribution in any form to anyone is expressly forbidden. Authors alone are responsible for the contents of their articles. The journal owns the copyright of the articles. The publisher shall not be liable for any loss, actions, claims, proceedings, demand, or costs or damages whatsoever or howsoever caused arising directly or indirectly in connection with or arising out of the use of the research material. All authors are requested to disclose any actual or potential conflict of interest including any financial, personal or other relationships with other people or organizations regarding the submitted work.

