Education
December 10, 2025
Language model selection for Enem essay scoring
Heitor Dutra de Assumpção; Adriana Camargo de Brito
Summary prepared by the ResumeAI tool, an artificial intelligence solution developed by Instituto Pecege focused on synthesis and writing.
This study verifies the feasibility of using Language Models (LLMs) as automatic evaluators, under the “LLM-as-a-Judge” paradigm, for the correction of essays from the National High School Exam (ENEM). The research systematically compares the evaluation quality, consistency, latency, and computational cost of five distinct models, subjected to a reproducible protocol. A single and stable prompt was employed, designed to instruct the models to strictly apply the official exam rubric, unfolded into the five competencies (C1 to C5). The evaluation was conducted on a public corpus of essays with a maximum score (1,000 points) in the 2022 to 2024 editions of ENEM, using an automated system for standardized collection of performance metrics and output compliance validations.
The grading of essays in large-scale assessments like ENEM represents a significant logistical and financial challenge. In Brazil, the traditional procedure, involving pairs of human graders and audits, incurs high operational costs and prolongs the time for result dissemination. Estimates point to direct costs of millions of reais associated with human grading (Oliveira Júnior, 2025), driving the search for technological solutions. The challenge is to develop systems that preserve adherence to the complex official rubric and introduce gains in scale, speed, traceability, and auditability in each judgment.
The literature on Automated Essay Scoring (AES) has a long history. Its origins date back to the Project Essay Grade (PEG), which used superficial predictors like word length to estimate a score (Page, 1966). More sophisticated systems, such as e-rater, have incorporated richer linguistic indicators and demonstrated high agreement with human graders (Attali; Burstein, 2006). Despite advances, the field of AES faces criticism regarding a lack of transparency, potential negative pedagogical effects, and the risk that systems may be susceptible to gaming strategies by students (Dikli, 2006; Perelman, 2014).
In the ENEM, the grading is anchored in five detailed competencies, requiring a multifaceted evaluation. A practical limitation for AAR research in Brazil is the data policy of the National Institute for Educational Studies and Research Anísio Teixeira (INEP), which publicly releases only essays with a maximum score of 1,000 points. This restriction prevents the application of traditional supervised learning approaches, which require a diverse dataset. Given this, the present study adopts a methodological design that prioritizes protocol transparency, direct comparability between models, and the implementation of rigorous automatic controls for validating the format and arithmetic of the outputs.
The “LLM-as-a-Judge” approach is an alternative aligned with recent studies exploring the capacity of LLMs to perform complex judgments based on explicit instructions (Liu et al., 2023; Zheng et al., 2023). The strategy consists of instructing the LLMs, through a detailed prompt, to apply the ENEM rubric and return the evaluation in a controlled JSON format, containing discrete scores for each competency and the final sum. This formulation allows for a direct and fair comparison between models under identical criteria. The proposed system materializes this approach in an automated pipeline in Python, which manages text reading, prompt assembly with the theme of each year, execution of multiple repetitions per essay for statistical robustness, and standardized metric collection.
The experimental pipeline was implemented in Python, automating the workflow. The system reads each essay, inserts the correct topic into the prompt, and submits the text to each of the five LLMs. To evaluate stability, each essay was processed five times by each LLM, with an inference temperature of 0.0 to maximize determinism. On each run, the system recorded detailed metrics, including input, output, and total tokens, as well as response and processing times. Data analysis combines descriptive statistics with indicators such as the sum error rate and hit count. The choice of models covered a diversity of architectures, including dense models like Llama 3.3, Mixture of Experts (MoE) models like Llama 4, and a composite system like Groq/Compound. Architectural differences, such as attention mechanisms (VASWANI et al., 2017; SHAZEER, 2019; AINSLIE et al., 2023) and post-training type (OUYANG et al., 2022), impact the balance between accuracy, latency, and cost.
The limited availability of public corpora for AES in Brazilian Portuguese is a historical obstacle. Initiatives like Essay-BR represented an advance by offering annotated essays (Marinho et al., 2021; Marinho et al., 2022), but the origin of the texts, from simulators, may introduce biases. Recent efforts have focused on improving data curation and documentation (Silveira et al., 2024), but the lack of broad and representative samples persists. The methodology of this work partially bypasses this limitation by focusing on a protocol that does not depend on fine-tuning, but on the models’ ability to follow complex instructions.
The evaluation metrics were selected for a multidimensional view of performance. In addition to descriptive statistics for scores, tokens, and latency, an analysis of output compliance was performed, verifying the validity of the JSON format and the correctness of the arithmetic sum. Each model’s ability to hit the reference score (1000 points) was quantified, as was the proportion of scores in close ranges (above 900 and 800). For latency, the Empirical Cumulative Distribution Function (ECDF) was used, a visual tool to compare the distribution of response times between models. This analysis allows for the identification not only of the most accurate model but also for understanding the operational trade-offs for large-scale implementation.
The results of the output validation revealed significant operational differences. The consistency between the final grade and the sum of competencies was the first metric. The Llama 4 model presented inferior performance, with a sum error rate of approximately 45.6%, indicating a systematic failure to follow a basic prompt instruction. In contrast, Llama 3.3 demonstrated perfect consistency, with 0% sum errors. The other models presented negligible or null error rates. To ensure the integrity of the analyses, the “aggregated grade” was recalculated for all runs as the sum of the five competencies, replacing the LLM’s “final grade” when there was a discrepancy.
In aggregate performance, Llama 3.3 emerged as the winner. It hit the 1000-point benchmark score in 85% of runs, a result superior to competitors, and exhibited the highest average and median aggregate scores, with the lowest standard deviation, indicating high consistency. The Groq system (Orchestrated) came in second place in quality. GPT OSS showed intermediate performance, with greater variability. Llama 4, in addition to operational issues, obtained lower scores, while Gemma 2 presented the weakest performance, with no evaluation reaching the 800, 900, or 1000 point ranges. Llama 3.3’s superior performance can be attributed to its dense architecture and large number of parameters.
The analysis by individual competence showed that, for almost all models, Competence II (Understanding of the proposal and genre) received the highest average scores, suggesting that this task is more accessible for LLMs. The exception was Gemma 2, which assigned its highest average score to Competence I (Mastery of standard language). The Llama 3.3 and Groq (Orchestrated) models demonstrated coherence in their evaluations across competences. The interpretation of these results is limited by the nature of the corpus, which reduces discriminability. The analysis by year/theme indicated that the 2023 theme (“Challenges for addressing the invisibility of care work performed by women in Brazil”) seemed to present slightly greater difficulty, but the performance ranking among LLMs remained consistent.
The computational cost and latency metrics provided crucial insights. Token analysis showed that the Groq (Orchestrated) system used the most tokens, due to its extensive prompt. The Llama family models and Gemma 2 were more efficient. This difference correlated with latency. Groq (Orchestrated) and GPT OSS were the slowest, with response times of up to 30 seconds and 5 seconds, respectively. The fastest group, composed of Gemma 2, Llama 4, and Llama 3.3, showed latencies consistently below 0.5 seconds. Within this group, Gemma 2 was the fastest, followed by Llama 4 and Llama 3.3. The analysis highlights a clear trade-off between quality, cost, and speed. Llama 3.3, the model with the best accuracy, was not the fastest, but operated within a low latency range, making it viable for large-scale applications.
The study has important limitations. The main one is the restriction of the corpus to 1000-point essays, which prevents the evaluation of the models’ behavior on lower scores and may overestimate accuracy. Future work should seek access to more diverse corpora. Operational errors, such as the high sum error rate of Llama 4, highlight the need to implement validation and post-processing layers. The variation in cost and latency between architectures (dense vs. MoE vs. orchestrated) is a critical factor for scalability. Issues regarding generalization, biases, and contamination by public data require continuous investigation.
In summary, among the evaluated models, Llama 3.3 demonstrated the best overall performance, combining high accuracy, operational consistency (0% sum error), and low standard deviation. The Groq system (Orchestrated) positioned itself as a high-quality alternative, but with a penalty in token consumption and latency. GPT OSS presented intermediate results with greater variability, while Llama 4 proved operationally flawed and Gemma 2, although faster, was the least accurate. The recommendations for institutional implementation are: Llama 3.3 should be prioritized when quality and consistency are critical; Gemma 2 can be an option when minimum latency is the main requirement; and the Groq system (Orchestrated) is only justified if its orchestration capabilities compensate for the high computational cost.
The study provides a framework for selecting LLMs in automatic evaluation tasks, basing the decision on the balance between quality, speed, and cost. The research validates the “LLM-as-a-Judge” approach as a promising methodology, but highlights the need for rigorous protocols, continuous validations, and an understanding of each model’s limitations. It is concluded that the objective was achieved: the feasibility of using LLMs as evaluators was demonstrated, with the Llama 3.3 model presenting the best balance between accuracy, consistency, and computational cost for the task of correcting ENEM essays.
References:
AINSLIE, Joshua; LEE-THORP, James; DE JONG, Michiel; ZEMLYANSKIY, Yury; LEBRÓN, Federico; SANGHAI, Sumit. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023.
ATTALI, Y.; BURSTEIN, J. 2006. Automated essay scoring with e-rater® v.2. The Journal of Technology, Learning, and Assessment, 4(3): 1–29.
DIKLI, S. 2006. An overview of automated scoring of essays. The Journal of Technology, Learning, and Assessment, 5(1): 1–36.
GOOGLE. Gemma 2-9B-IT — Model Card. 2025. Available at: https://console. groq. com/docs/model/gemma2-9b-it and https://huggingface. co/google/gemma-2-9b-it. Accessed on: Sep. 26, 2025.
GROQ. Compound Systems — Overview. 2025. Available at: https://console. groq. com/docs/compound and https://console. groq. com/docs/compound/systems. Accessed on: Sep. 26, 2025.
HOLTZMAN, A.; BUYS, J.; DU, L.; FORBES, M.; CHOI, Y. 2020. The curious case of neural text degeneration. In: International Conference on Learning Representations (ICLR), 2020, Addis Ababa, Ethiopia. Available at: https://openreview. net/forum? id=rygGQyrFvH. Accessed on: Jun. 9, 2025.
INSTITUTO NACIONAL DE ESTUDOS E PESQUISAS EDUCACIONAIS ANÍSIO TEIXEIRA [INEP]. 2024. The Enem 2024 essay: participant’s booklet. Available at: https://download. inep. gov. br/publicacoes/institucionais/avaliacoeseexamesdaeducacaobasica/aredacaonoenem2024cartilhadoparticipante. pdf. Accessed on: Jul. 5, 2025.
LIU, Y.; et al. 2023. G-Eval: NLG evaluation using GPT-4 with better human alignment. In: Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023, Singapore. Proceedings… pp. 2511–2522. Association for Computational Linguistics.
MARINHO, J. C.; ANCHIÊTA, R. T.; MOURA, R. S. 2021. Essay-BR: a Brazilian corpus of essays. Available at: https://arxiv. org/abs/2105.09081. Accessed on: Jul. 1, 2025.
MARINHO, J. C.; ANCHIÊTA, R. T.; MOURA, R. S. 2022. Essay-BR: a Brazilian corpus to automatic essay scoring task. Journal of Information and Data Management, 13(1): 65–76.
META. Llama-3.3-70B — Model Card. 2024. Available at: https://www. llama. com/docs/model-cards-and-prompt-formats/llama3_3/ and https://huggingface. co/meta-llama/Llama-3.3-70B-Instruct. Accessed on: Sep. 26, 2025.
META. Llama-4-Scout-17B-16E — Model Card. 2025. Available at: https://huggingface. co/meta-llama/Llama-4-Scout-17B-16E and https://build. nvidia. com/meta/llama-4-scout-17b-16e-instruct/modelcard. Accessed on: Sep. 26, 2025.
NATIONAL COUNCIL OF TEACHERS OF ENGLISH [NCTE]. 2013. Machine Scoring Fails the Test: NCTE Position Statement on Machine Scoring. Available at: https://cdn. ncte. org/nctefiles/press/machinescoring-2013. pdf. Accessed on: Jul. 1, 2025.
OPENAI. gpt-oss-120b & gpt-oss-20b — Model Card. 2025. Available at: https://openai. com/index/gpt-oss-model-card/. Accessed on: Aug. 14, 2025
OUYANG, Long et al. Training language models to follow instructions with human feedback. In: Advances in Neural Information Processing Systems (NeurIPS), 2022.
PAGE, E. B. 1966. The imminence of grading essays by computer. Phi Delta Kappan, 47(5): 238–243.
PERELMAN, L. 2014. When “the state of the art” is counting words. Assessing Writing, 21: 104–111.
PONTES, A.; PRISCILLA, R.; LOPES, A. Chapter 24 Automatic essay scoring. [s. l: s. n.]. Available at: <https://brasileiraspln. com/livro-pln/2a-edicao/parte-aplicacoes/cap-aes/cap-aes. pdf>. Accessed on: Sep. 27, 2025.
SHAZEER, Noam. Fast Transformer Decoding: One Write-Head is All You Need. 2019.
SILVEIRA, I. C.; BARBOSA, A.; MAUÁ, D. D. 2024. A new benchmark for automatic essay scoring in Portuguese. In: International Conference on Computational Processing of Portuguese (PROPOR), 16., 2024, Santiago de Compostela, Spain. Proceedings… pp. 228–237. Association for Computational Linguistics.
VASWANI, Ashish et al. Attention Is All You Need. In: Advances in Neural Information Processing Systems (NeurIPS), 2017.
Executive summary from the Final Project of the Specialization in Data Science and Analytics from the MBA USP/Esalq
Learn more about the course; click here: