Working paper
Can Large Language Models deliver human-equivalent essay feedback without prompt-specific pre-training?
- Abstract:
- This study investigates the potential of general-purpose large language models (LLMs) to support essay assessment at scale without the bespoke tuning that currently characterises most AI applications. Using structured prompts and pre-existing analytic rubrics, the research examined whether LLMs could generate scores and formative feedback that align with human judgement in both technical and pedagogical dimensions. Drawing on a dataset of 17,000 student essays across four iterative phases, the study evaluated whether prompt engineering alone, without task-specific pretraining, can achieve levels of scoring accuracy and feedback quality comparable to those of trained educators. Findings demonstrate that, within defined conditions, general-purpose models can attain human-equivalent performance. The standard model improved from initial inconsistency to the level of an average human grader, while the advanced model achieved consistent expert-level agreement. These results were validated across multiple genres, grade levels, and rubric structures using psychometric indicators including Quadratic Weighted Kappa (QWK), Intraclass Correlation Coefficient (ICC), and Root Mean Squared Error (RMSE). Notably, the models performed strongly on both mechanical and higher-order traits, such as comprehension and elaboration. The study concludes that, when carefully designed and pedagogically contextualised, LLMs can enhance the scalability, consistency, and responsiveness of formative assessment without displacing human interpretive judgement.
- Publication status:
- Published
Actions
Authors
- Publisher:
- University of Oxford
- Place of publication:
- Oxford, UK
- DOI:
- Language:
-
English
- Keywords:
- Pubs id:
-
2464551
- Local pid:
-
pubs:2464551
- Deposit date:
-
2026-10-04
- ARK identifier:
If you are the owner of this record, you can report an update to it here: Report update to this record