The synthesis below is being revised and may be incomplete or out of date. Check back soon for the refreshed version.
Papers cited with a reference number in the synthesis below have been checked by a human expert. However, besides pointing to available literature, we do not evaluate the scientific quality of methodological rigor of studies. Please consult original publications and apply your own critical judgment.
Applications
Generative AI is being widely adopted across disciplines to substitute, augment, modify, and redefine various facets of feedback provision in education. At the substitution level, it primarily reduces teacher workload by supporting qualitative feedback delivery for assignments and exercises, while augmentation can enhance feedback quality through alignment with established pedagogical frameworks. Modification transforms feedback administration through automation, enabling immediate delivery and possible integration with learning management systems. Redefinition more profoundly reimagines feedback provision, encompassing innovative approaches such as feedback within roleplay scenarios and simulation-based learning environments.
Substitution: Generating Feedback
Given generative AI’s sophisticated linguistic capabilities, the most extensive application has been in providing feedback on written assignments, including for example essays, reports, and translation tasks, where AI systems offer linguistic corrections, structural improvements, and guidance on content quality and argumentation [153], [185], [38], [1], [10], [115], [16], [8], [129], [130], [235], [213], [15], [44], [173], [26], [102], [249], [176], [110][223], [80], [11], [20].
Feedback on programming assignments represents a natural extension of these linguistic abilities, as programming languages themselves consist of syntax, semantics, and structural rules. Generative AI can identify syntactic or logic errors, and provide feedback on code efficiency, style, readability and documentation [100], [101], [44], [137], [221], [145], [48], [70], [217], [9], [144].
Beyond purely linguistic feedback on natural or programming languages, AI can be employed to provide feedback that draws upon knowledge domains. This involves evaluating students’ understanding of subject matter, assessing the quality of their reasoning and problem-solving approaches, and their disciplinary performance [49], [32], [42], [66], [19], [208], [172], [77], [227].
Augmentation: Enhanced Feedback Design
Generative AI can enhance educational feedback in multiple ways: alignment with specific learning objectives, adherence to established pedagogical frameworks, application of evidence-based feedback principles, incorporation of domain-specific knowledge of errors, implementation of thoughtful design principles, and feedback delivery through diverse modalities and styles [205], [210], [35], [184], [139], [71], [153], [155], [99], [188], [192], [186], [58].
Modification: Automatic Integrated Feedback
The affordance of generative AI to generate adapted content on the fly enables modification of feedback delivery in several ways, including automation of feedback provision and it’s integration with learning platforms, and realtime or immediate delivery of feedback that is personalized or adapted to individual learner needs by representing a dynamic student state. These applications are prevalent in the computer science domain [152], [155], [151], [145], [219], [67], [142], [95], [72], [104], [79], [91], [93], [6], [186], [9], and in the educational field more generally [216], [135], [58], [239], [112]. Realtime delivery of feedback is particularly relevant for language learning, where students can get feedback on speaking performance during realtime language practice [143], [79], [18], [192], [122].
Redefinition: Interactive Feedback and Simulations
Finally, generative AI can redefine feedback integration by creating roleplay and simulation-based learning environments where feedback is either embedded within the interactive experience or delivered immediately afterward. These applications are most prevalent in the medical field [94], [24], [182], [167], [30], [236], [33], [29], [25], [162], with an additional example in physics [87] and teacher education [157].
Measures and Outcomes
Feedback Quality
The question of how AI-generated feedback compares to human feedback has attracted substantial empirical attention, and the evidence divides broadly into two camps: studies documenting AI feedback as effective, comparable to, or in some respects superior to human feedback, and studies reporting mixed results, domain-specific shortcomings, or fundamental limitations in comparability.
Several studies have argued that AI-generated feedback can be an effective support tool [6], [44], [50], [184], [235] and is comparable to or exceeds human feedback in some cases [34], [173], [20], [155], [42], [210], [53], [185], [172], [66].
However, other studies have reported mixed results, noting that AI-generated feedback can be ineffective or less effective in some areas, or is inconsistent and unreliable [100], [80], [144], [16], [157], [151], [55] and/or low comparability with human feedback [85], [249], [179], [160], [112], [23], [48], [137].
Feedback Perception
Beyond the objective quality of AI-generated feedback, students’ subjective experience of this feedback represents a crucial factor in determining its ultimate educational impact. Students’ subjective perception of AI-generated feedback is shaped by a constellation of factors, including their prior attitudes toward AI technology and their experience with AI tools [165], [205], [4]. Often the preference for human or AI feedback depends on the student, with one study showing an even split of preference between students [110].
Perhaps surprisingly, students often cannot reliably distinguish between AI-generated and human-generated feedback, frequently rating AI feedback quite favorably when they are unaware of its source, likely because of its greater length and comprehensiveness [217], [66]. The picture changes substantially when the source of feedback is disclosed. Research reveals a complex and sometimes contradictory pattern of responses. On one hand, some studies indicate that students prefer human feedback and express skepticism toward AI-generated responses [197], [104]. On the other hand, natural, flexible interaction experiences, disclosure transparency and student feedback literacy are correlated with positive attitudes [225], [52], [244], [93].
These findings suggest that, beyond the quality of the feedback content, the flexibility of the interaction design is a critical determinant of how students perceive AI-generated feedback.
Engagement
Research on student engagement with AI-generated feedback reveals a multidimensional picture that spans cognitive, affective, and behavioral dimensions, with evidence accumulating on both the benefits and the limitations of this mode of feedback delivery.
Many studies report positive effects across cognitive, affective, and behavioral engagement, often crediting the interactive and immediate nature of AI feedback [47], [5], [218], [8], [251], [239], [135], [170], [87].
Enhanced engagement with feedback has been associated with subsequent revision behavior, representing one of the most concrete outcome measures in this literature [95], [175], [115], [139].
Despite these positive findings, some studies suggest that student engagement with feedback remains generally low regardless of source. One of the most striking examples of the sheer scale of non-engagement: an analysis of 655 writing tasks involving 14,236 students found that, even though AI-generated feedback led to higher revision frequencies compared to teacher feedback alone, approximately half of all students still did not revise a single character in their text after receiving generative AI feedback [176], [120]. This finding underscores that the availability of AI feedback does not automatically translate into behavioral engagement.
Other drawbacks are cognitive overload and the time-consuming nature of processing AI feedback [41], [170], [21].
Taken together, the evidence suggests that AI-generated feedback can meaningfully stimulate cognitive, affective, and behavioral engagement, particularly when it is immediate, personalized, and integrated into iterative revision workflows. However, substantial proportions of students do not act on AI feedback, while some experience cognitive overload or find the feedback insufficiently nuanced. Finally, it is important to note that excessive usage has been linked to decreased performance outcomes [152].
Educational Outcomes
The body of evidence on AI-generated feedback in education spans a wide range of domains and outcome types, and the findings, while broadly encouraging, are nuanced in important ways.
In writing and language learning contexts, AI feedback has been associated with improved skills and performance [223], [129], [81], [247], [97], [125], [213], [69], [46], [15], [233], [143], [58], [1].
Similar positive outcomes extend beyond language domains, with studies documenting enhanced learning performance and academic achievement across various educational disciplines [30], [107], [188], [170], [239], [91], [104], [110], [135].
These benefits are often mediated by increased student engagement [218], [5], [87].
Beyond domain-specific learning outcomes, AI-generated feedback has been linked to various psychological and metacognitive benefits. These include increased motivation [69], [188], [236], reduced anxiety [232], [61], enhanced autonomy and perceived competence or self-efficacy [251], [123], [51], and improvements in self-regulated learning [79], [91].
However, the evidence also contains cases where significant learning gains were absent [229], [152].
Limitations and Recommendations
Enhancing AI Feedback
A recurring diagnosis across the literature is that off-the-shelf large language models, when prompted without domain grounding or pedagogical alignment, tend to produce feedback that is generic, inconsistent, or insufficiently actionable for learners. Developers have responded by developing a range of technical strategies aimed at making AI feedback more precise, contextually appropriate, and educationally sound. Three broad families of approaches have emerged: retrieval-augmented generation, fine-tuning and reinforcement learning, and multi-step architectures.
Retrieval-arugmented generation (RAG) ensures that feedback is contextually appropriate by drawing from relevant educational resources and domain-specific knowledge bases[22], [192]. For example, a programming education system called Tutor Kai explicitly linked its feedback to lecture videos and course content through retrieval [122].
Fine-tuning and reinforcement learning represent a more resource-intensive but potentially more powerful route to aligning model outputs with pedagogical needs. LLMs can be finetuned using specialized datasets tailored to particular academic disciplines, or reinforcement learning techniques can guide models toward producing feedback that aligns with educator objectives [192], [210], [6]. For example, in the context of mathematical tutoring dialogue, fine-tuned smaller models trained on the MATHDIAL dataset significantly surpassed the performance of much larger prompted LLMs, while keeping the rate of directly telling students answers lower [148].
Multi-step architectures decompose the feedback generation process into distinct, specialised stages, typically separating evaluation from pedagogical guidance and adding a validation layer to check quality criteria before feedback reaches students. This modular approach can be implemented through various architectural strategies. Multi-component systems may combine LLM-based modules with traditional rule-based c omponents, capitalizing on the strengths of each approach. Alternatively, multi-agent frameworks can deploy several instances of a language model, each guided by distinct system prompts that define specific roles in the feedback pipeline [54], [71], [43], [77], [192]. For example, in programming assessment, one system incorporated separate stages for strict code evaluation, review quality maintenance, and filtering of unnecessary feedback [155].
The most pedagogically effective systems will combine several of these strategies: grounding outputs in domain knowledge through retrieval, aligning generation with educator objectives through fine-tuning, and decomposing the feedback task into specialised stages. Crucially, human oversight needs to be incorporated at critical validation points.
Designing Effective Human-AI Interaction
The accumulated evidence across disciplines converges on a foundational principle: generative AI functions most effectively in feedback contexts not as an autonomous replacement for human judgment but as a support tool operating within carefully designed human-AI workflows. Fully automated feedback systems, while scalable, carry documented risks that only hybrid architectures can adequately address.
Human-in-the-loop designs, in which AI generates a first pass that educators review before release, are particularly widely recommended as a practical bridge between the scalability of automated systems and pedagogical expertise, offering complementary strengths where teachers can for example address higher-order concerns while AI handles surface-level feedback [192], [115], [227].
Several interaction design principles emerge as critical to making human-AI feedback collaboration effective. Key principles include promoting transparency about AI limitations, maintaining instructor control, and creating natural, flexible interaction experiences [145], [225], [35], [84]. Systems must also leverage generative AI’s capabilities to promote inclusion and ensure equitable access for diverse learner populations [216], [252].
Preventing over-reliance on AI tools remains crucial to avoid diminishing critical thinking and creativity. Both educators and students require comprehensive training in AI literacy and feedback literacy to ensure responsible and effective use of these technologies. These efforts should operate within regulatory frameworks that prioritize responsible use and support ethically grounded pedagogical implementations rather than purely technical solutions [206], [154], [234].
System Implementation Recommendations
Successful deployment requires attention to system scalability and performance optimization to handle educational workloads effectively [221], [219], [58]. Furthermore, privacy compliance represents a critical concern, as educational data requires special protection measures [192]. To address privacy concerns while maintaining functionality, institutions should consider implementing open-source LLMs that allow for local deployment and greater data control [101], [70], [44].
Conclusion and Future Directions
Generative AI can transform educational feedback by enabling immediate feedback that is aligned with pedagogical principles, and supporting entirely new formats such as feedback within simulations and interactive feedback. Future work should focus on advancing feedback quality through refined approaches such as retrieval-augmented generation and multi-step systems, with consistent integration of human expertise as an essential component. Research should also emphasize AI’s potential for creating innovative feedback delivery methods that incorporate simulation and real-time interaction.
The synthesis below is being revised and may be incomplete or out of date. Check back soon for the refreshed version.
Applications
Generative AI applications in assessment evaluation span a wide range of intended innovation, and situating them within the ISAR model (Bauer et al., 2025) helps clarify both their promise and their risks before examining specific performance evidence. At the substitution level, AI systems are deployed to replace or assist human graders in routine scoring tasks, with the primary justification being efficiency rather than any qualitative improvement over existing practice — the threshold for success is simply adequate agreement with human judgment. Moving up the spectrum, augmentation describes systems designed to exceed the human baseline in some measurable respect, whether through greater scoring consistency across raters and contexts. At the furthest reach, redefinition encompasses evaluation modalities that are genuinely novel — forms of continuous, adaptive, or multimodal assessment that could not feasibly exist without generative AI capabilities. Cutting across all three levels, however, is the risk of inversion: AI grading systems that introduce bias, reward construct-irrelevant features of student responses, or produce systematic inaccuracies can undermine the validity and fairness of assessment regardless of how ambitiously they are conceived, making inversion a persistent concern that must be examined alongside evidence of intended gains.
Substitution: Replacing and Assisting Human Graders
The substitution literature spans a wide range of educational contexts and assessment types, united by a common design logic: AI systems are deployed to perform evaluation tasks that human graders would otherwise carry out, either taking over the process entirely or handling the bulk of scoring work while a human retains a supervisory or exception-handling role. The efficiency rationale is explicit across nearly all studies, even when quantitative time-savings data are sparse.
For example, in automated essay scoring: AESCORE: a lightweight, rubric-integrated automated essay scoring (AES) framework built with open-source lightweight LLMs that generates rubric-level predicted scores for EFL essays using engineered prompting strategies [202].
Table 1. Studies using generative AI for evaluation tasks, by type.
|
Type
|
References
|
|
Automated Scoring Essays, Argumentation, Writing
|
[202], [40], [179], [36], [217],
[187], [170], [54], [149], [165], [155], [135], [6], [173]
|
|
Computer Science and Engineering
|
</td>
|
|
Language and Literacy
|
</td>
|
|
Mathematics and Physics
|
</td>
|
|
Domain-General Education across Primary, Secondary, and Tertiary Settings
|
</td>
|
PAY ATTENTION TO DEPLOYED SYSTEMS
A customized GPT has scored student essays in music education [181]
Automated scoring of short answers, essays, and open-ended responses constitutes the most prevalent function in the corpus. Systems operating at this level ingest student-produced text and return numeric scores, sometimes with partial credit, across a variety of disciplines and education levels. In dentistry, a rubric-guided pipeline scored clinical short-answer responses across four domains [144], and open-ended handwritten responses via GPT-4o [34]. Scoring of SQL query responses has been evaluated using ChatGPT, Gemini, and Copilot against itemized marking rubrics [37]. Across general education contexts, benchmark-style comparisons of short-answer grading systems have been conducted using unified multi-dataset collections [132], and self-augmentation methods have been applied to BERT-based models for scoring free-text responses on small datasets [45]. A RAG-based system combining BM25 or FAISS retrieval with a small LLM generator has been used to score concept-map propositions [153]. A Telegram bot integrated with OpenAI GPT has scored quiz responses and provided immediate feedback [91]. A dissertation-grading application using ChatGPT has been compared against human evaluators and prior computational methods [180]. An AI-evaluator integrated into Moodle has automated checking of written work against teacher-defined criteria with configurable grading formats [122]. Claude Sonnet 4 and ChatGPT-4o have been deployed as automated evaluators using a standardized five-criterion rubric for comparative ranking [218]. A hybrid automated grading system for Chinese–Portuguese translation has implemented multi-dimensional translation quality assessment [159]. A GPT-4-based system has scored student writing on an EOP platform across rubric criteria [171]. . .
Short Text Factual Answer –> correctness
multilingual short responses [21].
An LLM-based automated short-answer scoring tool built on GPT-4 ingests rubrics and model answers, extracts answer components, allocates partial credit, and aggregates scores across multiple sampled runs to produce a final mark [129]. A locally deployed system using Ollama processes open-ended exam responses by retrieving relevant course content, generating an ideal answer, comparing student responses semantically, and assigning a numeric score; the same ninety-exam set that required approximately twenty-five hours of manual grading was processed in approximately forty minutes of computation, with instructor review of flagged cases reducing total human time to two to three hours, engineering, AUTOMATED & DEPLOYED [138]. see also [48]. A Bing Chat LLM has been used to score open-ended USMLE Step 1–style responses using a standardized rubric [210]. educational psychology, deployed ? [128].
n language arts, a web-based Flask application integrated with ChatGPT has generated rubrics from teacher input and applied them to student textual submissions [7]. SmartGrading is an open-source platform that accepts questions, sample solutions, and evaluation instructions and produces exportable numeric scores for large batches of open-ended responses, with described use cases including double-checking human grading and handling large coursework, customizable solution for automated scoring [55]. ChatGPT 4.5 was used in a pilot to grade final-exam open-ended responses in parallel with human lecturers using the same rubric, with the broader framework also supporting analytics dashboards and anomaly flagging [62]. CAMS, a locally deployable collaborative attention-based multilayer system, composes multiple small LLM modules for keyword extraction, semantic matching, and dynamic weighted scoring to produce numeric grades and structured reasoning for student answers [145].
ADD OCR to table, for handwritten texts ?
short text answer: [150]
Reflections: A multi-agent LLM pipeline has produced per-dimension ordinal rubric scores for learner reflections [211]. A GPT-4-based system has scored student reflections using a rubric and used those scores as features for grade prediction. A single-agent and multi-agent LLM system converts open-text weekly student reflections into numeric rubric scores for downstream at-risk identification and grade prediction [213]
reflections [146]
Pitches: An LLM-based pitch grading system fine-tuned on GPT-3.5 and GPT-4o-mini predicts rubric scores across five components of nursing students’ pitch scripts and generates personalized written feedback [161].
ChatGPT was configured to evaluate School Improvement Plans against a seven-dimension leadership-competency rubric on a five-point scale, with scores compared to human mentor and scholar evaluations [207].
A GPT-4-based system evaluated graduate student digital portfolios using a twenty-seven-item eComplex rubric, generating item-level scores with citations, dimension analyses, formative recommendations, and a global score across 120 portfolios, explicitly does not assign final grades [186].
ChatGPT 4.5 scored fifty-two lesson plans against a thirteen-item expert rubric under both structured and unstructured prompting conditions [152].
In computer science, a GPT-4o system has performed per-criterion quality evaluation of student-generated user stories against the QUS Framework [206]. A ChatGPT-based evaluator using the INVEST framework scored student-generated user stories on six criteria across three iterative sessions, aggregating results via median [214].
mathematic problem solutions, calculus submissions [43] mathematics [14].
physics problems & solutions [61]
AI raters (ChatGPT-5 and Gemini-Pro 1.5) have been configured with a five-point analytic rubric to evaluate C2-level Turkish translation texts [208],
science structured questions, designed for automation [195],
engineering assignment evaluation [142].
evaluate creativity, automated by design [10], Creativity: The Creativity Assessment Platform provides one-click automated scoring of responses to multiple creativity tasks using validated ML and transformer models, enabling batch scoring without programming or manual scoring effort [148]. Two automated creativity-assessment systems—an evidence-centred design computational pipeline and a ChatGPT-4o LLM scorer—evaluate student-created music programs across dimensions including divergent thinking, complexity, and expressiveness, applied to 383 student artifacts [182].
A customized in-house ML tool has transcribed OSCE audio and applied a transformer to assign item-level checklist scores [115]
SOAP notes: A GPT-4-based system called GRADES scored SOAP notes by submitting them via a Python script to the GPT-4o API with the complete rubric, producing structured JSON scores; the system graded eighty-three notes in nine minutes total compared with 415 minutes for faculty and 4,150 minutes for peer graders [158]. SOAP against 4 rubric categories: [178]
Code: correct programming exercise, e-assessment pipeline [3], evaluate programming assignments [46], [33], A multi-role AI agent system for programming education includes a dedicated Evaluator agent that assigns per-criterion rubric scores with rationales and aggregates a total score for submitted code [223]. GPT-4o was used with a structured rubric to grade programming assignments, with the study reporting that automated grading of 534 assignments took approximately 2.5 hours at a cost under $0.05 per assignment, compared with an estimated forty-five or more hours of human effort [169]. Five LLMs were evaluated as automated grading assistants for Python debugging problems using rubric-aligned prompts, serving as preliminary scoring and triage tools, evaluate bug fix performance [116]. BeGrading, a fine-tuned LLM trained on real and synthetic programming submissions, predicts numeric grades and evaluates code across criteria such as correctness, efficiency, and readability, with iterative re-evaluation when grade discrepancies exceed thresholds [16]. A Python-based framework integrating CodeBERT, TF-IDF, XGBoost, and DBSCAN performs rubric-aligned classification of code submissions alongside plagiarism detection and anomaly flagging [229]. coding assignment [2], computer science, automated code review, students directly exposed [5], [100], tested grading behaviors of different models [117], –> augmentation when scaling up availability of evaluation ?
hand-written short answers: An automated grading system integrating Google Cloud Vision OCR with a DeepSeek-R1 1.5B model scores handwritten student answers in independent and contextual evaluation modes [174].
An LLM-based automated grader deployed in a bioinformatics course uses rubric- and example-based prompting to return numeric scores, satisfied rubric criteria, and written feedback, with most submissions automatically graded and manual review requests occurring in only 0.6% of cases [136].
evaluate peer feedback quality [18]. [196]. A GPT-4-based post hoc scoring system using C.L.E.A.R. and C.R.E.A.T.E. frameworks has assigned rubric scores and produced structured feedback tables for peer-review submissions [168].
handwritten answers chemistry exams, includes handwriting (OCR) - Ethel - [188]
Clinical and medical education contexts have adopted substitution approaches for structured performance evaluation. An AI-based scoring application using GPT-o1 Pro and GPT-5 Pro automatically rates clinical interview transcripts against the twenty-five-item Master Interview Rating Scale, with per-case scoring time reduced by 58–68% relative to physician raters and throughput increasing from approximately six to eighteen cases per hour [123]. A GPT-4-based system applied to clinical history-taking dialogues automatically checks whether specified history-taking categories were addressed and returns structured JSON yes/no verdicts per category [1]. An LLM-based CPX practicing system produces per-item binary scores with rationales across nineteen items and aggregates a total score, with Claude averaging 43.2 seconds and GPT-4o averaging 50.6 seconds per case [26]. A proof-of-concept combining speech-to-text transcription with LLM-based evaluation automatically assesses emergency call dispatcher performance across multiple competency areas [227]. medical short answers: PASS, a specialized GPT-4-based system configured for postgraduate anesthesiology examination preparation, grades short-answer question responses on a zero-to-eight scale and generates examiner-style feedback across multiple independent iterations [184]. In medical short-answer grading, multiple LLMs produce scores and rationales, internally generated and expert-provided rubrics [175] assign numerical grades to students’ short open-ended answers in medicine [49].
Math: The MathEDU dataset and associated fine-tuned LLMs support automated answer correctness classification, error-step identification, and feedback generation for student mathematical problem-solving processes [197].
short text responses, economics, humans not exposed [48],
and the AutoSCORE framework uses a rubric-component extraction agent followed by a scoring agent to assign final scores, score student answers, designed for automation [140]
prompt grading: A web-based LLM-driven module auto-grades student-written prompts across binary dimensions and returns immediate item-level feedback within a prompting-literacy practice workflow [194].
text answer language exam, includes OCR for handrwiting, deployed? [176],
A GPT-4-based coherence scoring system rates discourse coherence of written test responses and generates rationales and revised model answers [13].
Augmentation: Enhancing Evaluation Quality and Feedback
Augmentation-level systems go beyond mere replacement of human graders by introducing qualitative improvements to the evaluation process itself — greater consistency, richer feedback, systematic pattern detection, or measurable gains in student learning. These enhancements are achieved through a variety of mechanisms that together distinguish augmentation from simple substitution.
A generative AI–driven personalized education assessment framework (five‑layer hierarchical architecture) that fine‑tunes an LLM (ChatGLM3‑6B) on 50,000 expert‑curated programming feedback instances to generate contextually responsive, diagnostic assessments and pedagogical feedback. The system integrates learner profiling and a domain knowledge graph to produce personalized assessment items, composite scores, interpretable reasoning chains, and corrective feedback in real time. [166].
A human-in-the-loop prototype generates MCQs from submitted student artefacts, deploys them via a VLE, and automatically compares student responses to AI-generated answers to produce grades, generating ten MCQs in under twenty seconds [134]. –> augmentation ?
Consistency and reliability. Several systems are designed primarily to reduce inter-rater variability and enforce more uniform rubric application than human evaluators typically achieve. The HARRS system combines a deterministic rule-based layer for structural compliance scoring with an LLM semantic layer that generates rubric-anchored qualitative feedback, explicitly mitigating the leniency bias observed in unconstrained LLMs and achieving statistical parity with human expert scores, hybrid pre-submission review tool (HARRS) [139]. SLASys, a BERT-based recommender integrated into Moodle, applies semantic comparison and fine-tuned classifiers to short-answer responses, with teachers reporting that successive classifier iterations produced increasingly unified evaluation criteria and reduced the inconsistency inherent in manual marking [114]. The AMTES system for medical history-taking employs parallelised sub-evaluations, original-text citation requirements, and multilevel post-processing verification to achieve low coefficient of variation and high intraclass correlation coefficients across repeated automated runs, demonstrating a form of internal consistency that human raters rarely sustain at scale [162]. In pharmacy education, a customised speech-to-text and transformer-based tool for OSCE checklist grading was reported to provide more objective and consistent scoring than faculty, who exhibited the variability typical of live assessment settings [115]. The AIvaluate conversational agent assists teachers during oral performance-based assessments by supplying LLM-generated structured prompts and real-time emotional state data, with qualitative reports from teachers suggesting that the system reduced perceived bias and standardised questioning across candidates [224].
Scoring rationales. The AERA framework distils ChatGPT-generated scoring rationales into a fine-tuned Long T5 model that simultaneously predicts scores and produces natural-language rationales; human evaluators judged AERA’s rationales to surpass ChatGPT’s own outputs on correctness of key-element matching and faithfulness to the rubric, and the process additionally surfaced label inconsistencies in the original human-annotated dataset [27]. AutoSCORE, a multi-agent LLM framework, first extracts rubric-relevant components from student responses into a structured JSON representation before passing them to a scoring agent, producing auditable evidence-to-judgment chains that make the basis of each grade transparent and debuggable [140]. An LLM-based grading system for short-answer questions in university courses used instructor-authored prompt templates with gold-answer exemplars to generate personalised written feedback at a class size that would otherwise have precluded individualised commentary, with students rating 74.7% of substantive feedback responses as helpful [120]. The DeepSeek-R1-based system for master’s students’ textual solutions generated structured rubric-based feedback tables with criterion-linked scores and textual recommendations; student survey data showed no significant difference from instructor feedback on most quality criteria, with only overall satisfaction significantly favouring instructors [203]. A GPT-4-based grading system for Jupyter notebook assignments in an astronomy course produced multi-sentence diagnostic feedback with detailed error descriptions and point-deduction rationales, contrasting sharply with the two-to-three-word comments typical of human graders under time pressure [215]. The study comparing GenAI natural-language grading explanations with BERT important-word highlights found that GPT-4-generated step-by-step rationales significantly improved the quality of human graders’ subsequent feedback — operationalised across three learner-centred dimensions — and were rated substantially higher on informativeness and comprehensibility than word-level attribution highlights, even though they did not accelerate the grading process [228]. Ethel, a course-grounded generative AI system for chemistry, provides formative feedback on handwritten work that identifies missing assumptions, units, and scale errors, and assists grading of open-ended exam responses against rubrics with confidence estimates, with students reportedly finding that challenging AI feedback actively supported their learning [188].
produce rationales coherence essay [13]
designed for automation, scores reflection, checks bias and consistency [211],
Pattern detection. Some augmentation systems derive their value from identifying systematic features across submissions at a scale impractical for human reviewers. The LLaMA 2-powered programming education platform captures execution results, error frequencies, and topic mastery patterns across submissions, feeding an analytics dashboard that surfaces systematic misconceptions at the cohort level rather than treating each submission in isolation [100].
Finally, a fully automated LLM-based formative testing system in tertiary computer science education generates test items, scores responses, identifies topical knowledge gaps at the individual level, and issues personalised study-material recommendations in a continuous loop — a modality in which the integration of item generation, automated scoring, gap detection, and individualised recommendation within a single adaptive cycle is what makes the evaluation form genuinely new rather than merely more efficient [53].
Video evaluation:
The zero-shot multimodal LLM pipeline for OSCE video segmentation uses frame-level classification and hidden Markov model decoding to identify physical examination periods within fifteen-minute recordings, reducing the video requiring human review by approximately 81% and enabling downstream pattern analysis of assessment-relevant behaviour at a scale that manual review cannot support [204].
LLMs for more nuanced evaluation of writing tasks, essay [147], add scoring of conversation transcripts here ? designed for automation, surgical non-technical skills, based on transcripts [185] OSCE procedural skills from video recordings [47] OSCE checklist from transcript: [222] root canal evaluation based on radio, A ChatGPT-4o system has scored root canal treatment radiographs across five predefined criteria [157]. A ChatGPT-based system has scored radiographic images of root canal treatments across specific radiological parameters [221] medical skills from conversation transcripts [121]
realtime essay scoring, designed for automation without human step, EsyGrade [167],
Redefinition: New Evaluation Modalities
Automated extraction of clinical experiences from free-text learning logs has been implemented using GPT-4-turbo to map extracted items to a curriculum template [124]. Automated scoring of classroom observation components has been implemented using multimodal models combining facial, speech, and transcript features to estimate encouragement and warmth ratings [12]. Automated evaluation of emergency call interactions has been implemented via speech-to-text transcription followed by LLM-based assessment of dispatcher performance [227]. Automated scoring of stereotypicality ratings has been performed by ChatGPT during development of a Similarity Rating Test for medical education [22]. Automated generation of marking schemes has been implemented using a hybrid multimodal Seq2Seq model for science structured questions [195], and ChatGPT has been used to generate marking memoranda for engineering tests [219]. Automated scoring of text complexity for Portuguese texts has been implemented using few-shot LLM prompting [183]. Automated scoring of error detection in experimentation protocols has been implemented using GPT-3.5 and GPT-4 to identify and label common student errors in secondary science inquiry tasks [11].
The most distinctive feature of redefinition-level applications is that the evaluation modality they instantiate could not exist without AI — not because AI makes it faster or more consistent, but because the modality itself depends on capabilities that have no viable human-only equivalent at the required scale, fidelity, or dimensionality.
The clearest cluster of such modalities involves simulated oral and clinical examinations. The Surgery Oral Examination LLM (SOE-LLM) is a fully automated Flask application built on GPT-4-turbo that simultaneously plays the role of oral examiner and evaluator: it conducts case-based surgical encounters drawn from MIMIC-IV clinical data, ingests the resulting transcript, and produces structured assessments of diagnostic reasoning and clinical decision-making across eight domains [102]. ECOSBot similarly deploys GPT-4o as both a standardized patient and a standardized examiner for nephrology OSCEs, performing fully automated binary checklist scoring across fourteen to eighteen items per station, generating competency-based global ratings, and providing real-time feedback — enabling on-demand OSCE practice that would be logistically impossible to staff with human standardized patients and examiners at equivalent frequency [163]. MedSimAI extends this paradigm by coupling an AI standardized patient engine with automated rubric-based scoring of clinical encounters, including the twenty-eight-item Medical Interview Rating Scale with quote-grounded justifications, binary checklist coding, competency dashboard aggregation, and proficiency-threshold detection, typically returning formative feedback within two to three minutes of encounter completion. A MedSimAI platform has produced rubric- and checklist-based assessments with quote-grounded justifications for clinical encounters [199].
Create evaluative standard: A related but distinct redefinition involves replacing the human expert panel that anchors scoring keys. The Script Concordance Test (SCT) is a format that requires a reference panel of clinical experts whose aggregate responses define partial credit; the fully automated system studied here uses six LLMs as that panel — authoring the vignettes, answering all thirty-two items, and deriving the scoring key from the modal AI response, with minority AI choices used to assign partial credit — making AI not merely a scorer but the constitutive source of the evaluative standard itself [216].
At an even higher dimensional scale, the MEDPI benchmark deploys a committee of LLM judges to score simulated patient–clinician conversations across one hundred and five dimensions mapped to ACGME competencies, with each judge producing category-level deliberation, dimension-level scores, evidence-linked rationales, and flags, aggregated across 7,097 simulated conversations — a form of multi-dimensional conversational evaluation that no human panel could perform at this volume or granularity, scored across 105 dimensions [198].
Beyond spoken clinical encounters, redefinition also appears in the assessment of physical artefacts that require computational perception. A PointNet-based neural network takes three-dimensional point-cloud representations of preclinical dental cavity preparations and predicts rubric-based scores directly from geometric structure, automating a form of evaluation that requires machine vision of three-dimensional form and has no practical human-only equivalent at scale [151]. In spoken language education, an IoT-integrated system combining automatic speech recognition with a Transformer-based pronunciation correction pipeline captures learner speech through sensors, performs phonetic error detection and classification, generates corrected phoneme sequences, and delivers real-time adaptive feedback — an ai-assisted modality in which the phoneme-level granularity and real-time correction of pronunciation errors across large learner populations is only achievable through the combination of ASR and neural sequence modelling [143].
medical interview transcripts, this could be counted as fully automated, redefinition ? [123].
clinical simulation evaluation [1] automated by design, clinical performance evaluation simulation [26],
VocalVerse — an efficient hybrid, full-song scoring architecture that leverages lightweight acoustic encoders and downstream modules to produce reference-free, multi-dimensional singing assessment (breath control, timbre quality, emotional expression, vocal technique) and descriptive textual feedback; supported by the Sing-MD dataset and evaluated with the H-TPR perceptual ranking benchmark. [212].
Evaluation Role and Function
The integration of generative AI into educational assessment spans a broad functional spectrum, ranging from fully automated systems that independently score, classify, and provide feedback on student work, to hybrid architectures in which AI serves primarily as a support layer that enhances the capacity and consistency of human evaluators. This structural distinction is not merely technical; it carries significant implications for questions of accountability, interpretability, and pedagogical validity, since a system operating autonomously bears different epistemic and ethical responsibilities than one whose outputs are filtered through human judgment. Understanding where a given application falls on this spectrum is therefore essential for accurately interpreting its capabilities, recognizing its limitations, and determining the conditions under which its use is appropriate in educational contexts.
AI Role: Full Automation, Assistance, and Hybrid Approaches
The structural positioning of AI relative to human evaluators varies considerably across the literature, spanning three broad configurations: systems that operate as the sole and final decision-maker (fully automated), systems that produce outputs a human then reviews or overrides (AI-assisted), and arrangements where the boundary between AI and human judgment is ambiguous or variable (in-between or hybrid).
Fully automated systems
BY DESIGN THESE SYSTEMS DO NOT PLAN FOR A HUMAN REVIEW STEP BUT OFTEN THIS IS BECAUSE THEY WANT TO COMPARE PERFORMANCE WITH HUMAN SCORING DOESN’T MEAN THEY ARE DEPLOYED THAT WAY
MED simulations and CS automated code review are examples of natively automated setups some other studies are clearly designing the system to be automated/realtime, without human review step
constitute the largest single group and appear across virtually every educational domain and level. In these configurations the AI produces a score, grade, or evaluative judgment that is treated as the output of record, without a mandatory human review step. At the tertiary level, fully automated graders have been deployed for written assessments in medicine [54], [150], [178],
Not actually deployed in an actual student setting, mathematic problem solutions [43], Same, physics problems & solutions [61],
computer science, automated code review, students directly exposed [5], [100], tested grading behaviors of different models [117],
language arts, human subjects not exposed [40], [173], LLMs for more nuanced evaluation of writing tasks [147], essays, human subjects exposed [155], text answer language exam, includes OCR for handrwiting [176],
short text responses, economics, humans not exposed [48],
reflections [146], medical skills from conversation transcripts [121], medical interview transcripts, this could be counted as fully automated, redefinition ? [123]. clinical simulation evaluation [1] automated by design, clinical performance evaluation simulation [26], automated by design, clinical case study simulation and evaluation [102] simulate patient, include evaliation, automated [163],
essays, not sure for fully automated grading [149],
realtime essay scoring, designed for automation without human step, EsyGrade [167],
science structured questions, designed for automation [195],
designed for automation, creativity assessment[182]
and multilingual short-answer items, designed for automation [21].
designed for automation, surgical non-technical skills, based on transcripts [185],
Within the fully automated category, most systems operate unconditionally — that is, every submission receives an AI-generated score without any threshold-based routing to human review. This is the dominant pattern for rubric-guided LLM graders applied to short-answer questions, designed for automation, export numeric scores without human override step [55],
evaluate creativity, automated by design [10], [148].
designed for automation [129],
designed for automation [211],
and the AutoSCORE framework uses a rubric-component extraction agent followed by a scoring agent to assign final scores, score student answers, designed for automation [140].
AI-assisted systems position the AI as a tool that produces scores, suggestions, or flagged outputs that a human evaluator then reviews, accepts, modifies, or overrides.
In medicine, flag discrepancies for human review [9], [172], [31].
hybrid pre-submission review tool (HARRS) [139]
In OSCE and clinical skills assessment, AI-assisted tools produce checklist-based scores and feedback that human examiners can review [222],
evaluate writing assignments, AI as a preliminary round grading (formative), followed by human-graded rounds [171],
The SLASys system in Moodle recommends rubric levels and provides probability scores for short-answer responses, explicitly positioning itself as a recommendation tool for teachers [114].
grading assistance for handwritten (OCR) responses in chemistry (ethel), explicitly flagging uncertainty for human review. If students do not agree with the AI decisions, the plan is to simply let them veto the judgment on particular rubric items, in which case a human will make the decision. Our experiments consistently showed that the system has a very low rate of false positives, so we do not risk artificial grade inflation through artificial intelligence - much further in deployment than any other study (pioneer) [188].
evaluation of essays, human-in-the-loop [187], evaluation of portfolios and the GPT-eComplex Assistant generates a scoring table with citations and formative recommendations but explicitly does not assign final grades [186].
A GenAI-powered grader produces step-by-step assessment explanations and integer scores in structured JSON to support human graders, alongside BERT-based attribution highlights [228].
A GPT-4–based tool assists educators in reviewing and improving assessments and feedback rather than grading student work directly [60].
Across all three role types, the literature reveals that the choice of automation level is shaped by domain, task type, and institutional context rather than by a principled framework. Fully automated systems predominate where tasks are well-defined and rubrics are explicit, such as short-answer scoring, code evaluation, and multiple-choice assessment. AI-assisted configurations are more common where the stakes of error are high, where outputs require professional judgment, or where the AI is explicitly positioned as a support tool rather than a replacement. Hybrid arrangements frequently arise not from deliberate design but from the absence of a clearly specified human review protocol, leaving the effective degree of automation to vary by implementation.
Evaluation Functions
The evaluation functions performed by AI systems across the literature span a wide and heterogeneous range of tasks, from assigning numeric scores to generating rich qualitative feedback, detecting errors, and triaging submissions for human attention. These functions often co-occur within a single system, making clean categorical separation difficult, but several dominant clusters emerge.
consistency support — where AI functions as a second rater to flag discrepancies or reduce grader load — appears across several studies, though it is often embedded within broader scoring workflows rather than constituting a standalone function. The SmartGrading application explicitly describes double-checking human grading as a use case [55]. A GPT-4-based system for calculus grading has conducted consistency checks including repetitive identification tasks [43]. A ChatGPT-based system for programming assignments has performed repeated-assessment consistency checks across multiple chat sessions [46]. A GPT-4-based system for SOAP note grading has used post-grading analytics to identify rubric items with high variability [158].
Flagging and triage — identifying submissions requiring special attention — appears as a distinct function in several studies. A fine-tuned RadBERT model has been deployed to detect discrepancies between trainee and attending radiology reports, triaging discrepant cases for educational use [9]. A prompt-tuned GPT-4 model has automatically identified missed diagnoses in trainee radiology reports and generated discrepancy feedback for human review [31]. A HIPAA-compliant GPT-4o system has detected predefined error types in resident reports and summarized discrepancies across a large corpus [172]. A SmartGrading application has described triaging and handling large coursework as an explicit use case [55]. A locally deployed LLM-based formative assessment system has implemented sequential tolerance-based triage, automatically accepting high-confidence scores, re-evaluating borderline cases, and flagging responses for mandatory human review [138]. A human-in-the-loop AI framework has generated MCQs from student work and deployed automated comparison of student responses to flag grading outputs for expert validation [134]. A ChatGPT-based system for scoring student writing has described detection and flagging of potential plagiarism as a draft-stage triage function [171]. A GPT-4-based post hoc scoring system has been tested on simulated edge cases including suspected plagiarism [168]. An LLM-based grading system for Jupyter notebooks has piloted offline LLM video analysis to flag potential academic-integrity intervals for human review [215]. An AI text-detection function has been described for identifying AI-generated student writing and informing grade adjustments [193].
A multi-agent LLM pipeline has included an Equity Monitor agent for bias and fairness checking of evaluator language [211].
Feedback generation — producing textual feedback alongside or instead of a score — is among the most widely distributed functions in the corpus and frequently co-occurs with scoring. In language arts, a GPT-4-based system has generated formative feedback across six writing dimensions for English as a non-native language students [4], and a ChatGPT system augmented with OCR has generated individualised feedback and cohort-level error summaries for handwritten Japanese essays [170]. A GPT-4-based system has generated corrective and constructive feedback for A1-level Spanish writing [173]. A ChatGPT-4 system has identified problems and generated assessment feedback across four L2 writing dimensions [217]. In medicine, an AMTES system has generated structured, evidence-cited feedback for medical history-taking dialogs [162], and a language-model-powered simulated patient system has returned structured JSON feedback on history-taking categories [1]. A GPT-4-based system has provided feed-up and feed-forward recommendations for medical MCQs using chain-of-thought prompting [44]. A ChatGPT-based system has generated detailed feedback for dental short-answer responses [144]. A GPT-4-based system has generated coaching-style feedback for non-technical skills in surgical transcripts [185]. In computer science, an AI-aided e-assessment pipeline has generated tailored feedback including hints for correction and code-style suggestions [3], and a genAI-powered automated code review workflow has generated checklist assessments and documented review results [5]. A GPT-4-based system has generated natural-language, step-by-step grading explanations as auxiliary information for human graders [228]. A web-based LLM-driven prompting-literacy module has delivered immediate, item-level binary feedback on student-written prompts [194]. An LLM-based automated grader for bioinformatics assignments has listed satisfied rubric criteria and generated written feedback [136]. In engineering, a GPT-4-based system has generated personalized feedback for circuit-analysis homework [142]. A GPT-4-based system has generated formative feedback for SOAP notes with post-grading analytics [158]. A GPT-4-based system has generated detailed feedback for political science short-answer responses [120]. A GPT-4-based system has generated formative feedback for student reflections in a computer science course [211]. An LLM-based CPX practicing chatbot has generated structured, specific feedback aligned with scoring criteria [26]. A GPT-4-based system has generated rationales for coherence scores and revised model answers [13]. A GPT-4-based system has generated explanatory feedback for medical MCQs [156]. A GPT-4-based system has generated feedback for nursing students’ pitch scripts [161]. A GPT-4-based system has generated feedback for student writing in a business course [186]. A GPT-4-based system has generated feedback for student reflective essays [155]. A GPT-4-based system has generated feedback for student essays in a macroeconomics course [48]. A GPT-4-based system has generated feedback for student essays in a language arts course [20]. A GPT-4-based system has generated feedback for student essays in a general education course [122]. A GPT-4-based system has generated feedback for student essays in a computer science course [2].
Accuracy and Agreement
The literature on AI evaluation accuracy and agreement is methodologically diverse, drawing on a wide array of metrics that reflect both the variety of evaluation tasks studied and the absence of a single dominant standard. Agreement with human grading, correlation coefficients, error metrics, classification statistics, reliability indices, and psychometric evidence each appear across subsets of the literature, and the choice of metric is often as revealing as the findings themselves.
The ASAG2024 benchmark evaluates automated short-answer grading systems across multiple datasets [132].
Agreement with Human Grading: Agreement Rates, Kappa, and Quadratic Weighted Kappa
Exact and adjacent agreement rates are among the most intuitively interpretable metrics and appear in a substantial portion of the literature, though they are rarely the sole measure reported. In automated essay scoring and short-answer grading contexts, exact agreement rates vary enormously depending on task complexity, rubric design, and model capability. Studies reporting strong agreement include one evaluating AI scoring of OSCE analytical checklists, where exact-match accuracy reached 96.92% and 94.98% for two stations respectively, exceeding faculty accuracy on one station [115]. Similarly, an automated scoring system for history-taking dialogues achieved item-level consistency rates of 95.75%–97.13% across three clinical cases [162]. At the other extreme, a study of AI grading of open-ended dental radiographic assessments found that the AI aligned with the human tutor’s standard only approximately 50% of the time [25], and a study evaluating ChatGPT on short-answer questions in health professions education reported that 63% of student answers had unacceptably high discrepancies between human and AI raters [150].
Cohen’s kappa and its variants are widely used to correct for chance agreement. Values reported across the literature span an enormous range. At the high end, a study scoring NAEP mathematics constructed-response items with fine-tuned transformer models reported a mean Quadratic Weighted Kappa (QWK) of 0.945 on the test set, approaching the human-to-human benchmark of 0.965, and meeting a pre-specified criterion of QWK difference less than 0.05 for nine of ten items [14]. A DistilBERT-based essay scoring system achieved QWK of 0.90 after three rounds of fine-tuning, substantially outperforming earlier baselines such as EASE (QWK = 0.69) [40]. For discourse coherence scoring using GPT-4, QWK reached 0.81–0.82 depending on prompting order, compared to a baseline of 0.39 [13]. An ASAP-based short-answer scoring system using GPT-4o achieved a mean weighted RMSE of 0.27, the best among tested models, though the primary metric in that study was error-based rather than kappa [132].
In contrast, many studies report kappa values in the poor-to-moderate range. A study comparing LLM scoring of medical short-answer questions found kappa values ranging from near zero to 0.61 depending on the model and question, with substantial variability across attempts [175]. A study of AI grading of endodontic radiographic assessments reported Cohen’s kappa values consistently below 0.2 [157]. For scoring of lesson plans, structured prompting yielded an ICC of 0.708 (interpreted as 70.8% agreement), while unstructured prompting produced an ICC of only 0.076 [152]. A study comparing LLM scoring of writing prompts to human raters found ICC values of 0.22 for ChatGPT and 0.05 for Google Bard, both described as poor and non-significant [19]. A particularly striking negative result came from a study evaluating ChatGPT on open-ended health professions questions, which reported Cohen’s kappa of −0.0786, indicating systematic disagreement rather than chance-level agreement [150].
The pattern across the kappa literature is clear: agreement is strongest when rubrics are highly structured, scoring scales are narrow, and tasks involve relatively unambiguous criteria such as factual recall or checklist-based clinical skills. Agreement degrades substantially for tasks requiring interpretive judgment, nuanced reasoning, or evaluation of higher-order constructs such as reflection quality, clinical reasoning, or creative writing. A study of reflective writing scoring found that AI–human ICC was 0.71 overall but fell to 0.52 for the “feelings” dimension and 0.59 for “analysis,” while remaining stronger for surface-level dimensions such as clarity and language mechanics [146]. Similarly, a study of non-technical skills scoring from OSCE transcripts found that none of five domains reached the pre-specified AC2 threshold of 0.60, with values ranging from −0.153 to 0.297 [121].
A recurring finding is that the provision of structured rubrics substantially improves agreement. One study of open-ended handwritten student responses found that rubric-based prompting (Scenario 2) produced a Cohen’s kappa of 0.79 and a correlation of 0.89 with expert scores, compared to kappa of 0.09 and correlation of 0.12 without rubric guidance (Scenario 1) [34]. A study of essay scoring in a language learning context found that a customized GPT with teacher-provided examples achieved Pearson correlations of 0.961 (restricted essay) and 0.800 (extended essay) with human scores, compared to 0.070 and 0.580 respectively without examples [181].
A notable methodological issue is the prevalence paradox affecting kappa: several studies note that kappa values become unreliable or uninterpretable when one category is very rare or very common, and some adopt Gwet’s AC1 or AC2 as alternatives [doi:10.1016/j.caeai.2023.100177; doi:10.7759/cureus.100508; doi:10.3390/healthcare13233121]. The use of Fleiss’ kappa for multi-rater designs is also common in clinical and OSCE contexts [doi:10.1067/j.cpradiol.2024.08.003; doi:10.2196/77580; doi:10.3389/feduc.2026.1729156].
Correlation with Human Scores: Pearson and Spearman
Pearson and Spearman correlations are among the most frequently reported metrics, particularly in studies of essay scoring, short-answer grading, and clinical performance assessment. High correlations are reported in several well-resourced studies. A system for scoring virtual patient interview transcripts using the MIRS instrument achieved Pearson r = 0.90 (95% CI 0.78–0.96) against human raters, with Lin’s concordance correlation coefficient of 0.88 [123]. A study of automated scoring of clinical short-answer responses found Pearson r = 0.872 for DeepSeek-3 versus human scores, compared to r = 0.64 for ChatGPT-4 [144]. A system for scoring reflective essays using a hierarchical rubric-anchored approach (HARRS) achieved Pearson r = 0.98 with human consensus, compared to r = 0.46 for an unguided LLM baseline [139]. For automated scoring of physics solutions, GPT-4 with a mark scheme provided achieved Pearson r ≈ 0.80, the highest among tested models, while blind grading without a mark scheme yielded r ≈ 0.56–0.64 depending on the model [61].
Moderate correlations are common in more challenging tasks. A study of LLM scoring of calculus homework found correlations ranging from 0.19 to 0.82 across exercises and models [43]. For scoring of summaries against analytic rubric criteria, Pearson correlations between Claude and human raters ranged from r = .12 (Paraphrasing, non-significant) to r = .69 (Integration) [147]. A study of LLM scoring of programming assignments found Pearson correlations above 0.86 across three languages, but with Bland–Altman limits of agreement spanning approximately ±18 points on a 100-point scale for some languages [169]. For scoring of student reflections in a creativity assessment context, GPT-4o achieved r = 0.69 with human CAT ratings, compared to r = 0.48 for an automated ECD pipeline [182].
Weak or non-significant correlations are reported in several clinical and complex-judgment contexts. A study of AI scoring of OSCE performance found Pearson correlations between AI and human scores ranging from r = 0.19 to r = 0.26 across five criteria, none reaching statistical significance [157]. A study of ChatGPT scoring of OSCE domains found ICCs near zero or negative across all six domains [222]. For scoring of writing prompts by cognitive complexity, ChatGPT achieved ICC = 0.22 and Bard ICC = 0.05 with human raters [19]. A study of AI grading of SOAP notes found that restricted cubic spline modeling explained less than 10% of variance in human scores (R² = 0.01–0.098) [178].
Error Metrics: MAE and RMSE
Mean Absolute Error and Root Mean Squared Error are used primarily in studies treating grading as a regression problem, and their interpretation depends critically on the scale of the rubric. On a 100-point scale, a study of open-ended exam grading found MAE of 2.93 and RMSE of 3.46, which the authors interpreted as approximately one grade level [62]. A study of automated scoring of pitch scripts found overall RMSE of 2.81 and MAE of 2.24 on a 25-point scale, with component-level RMSEs ranging from 0.66 to 1.31 [161]. For scoring of dental cavity preparations from 3D point clouds, MAE was 0.82 on a 20-point rubric in training and 0.97 on the test set, with 50% of test-set predictions within ±1 point [151].
On shorter scales, a study of concept-map scoring (0–3 scale) found MAE of 0.973 and RMSE of 1.321 for the best-performing configuration [153]. A study of debugging problem scoring found Gemini achieved the lowest MAE of 0.724 and RMSE of 0.916 among five LLMs, while ChatGPT had the highest MAE of 1.750 [116]. For scoring of student reflections on a 0–3 scale, overall mean MAE was 0.467 [211]. A study of handwritten answer scoring found DeepSeek achieved MAE of 0.058 and RMSE of 0.147 on good-handwriting items, compared to BERT’s MAE of 0.3899 on poor-handwriting items [174], illustrating how preprocessing quality can substantially affect error metrics.
Bland–Altman analysis, which quantifies systematic bias and limits of agreement, is increasingly used alongside or instead of MAE/RMSE in clinical and health professions contexts. A study of AI scoring of OSCE skills found Bland–Altman biases of −6.66 and −4.59 for two AI models on a square-knot task, with auditory criteria showing larger negative biases [47]. A study of virtual patient interview scoring found a mean bias of +0.43 (SD 2.70) for the better-performing AI model, with 95% limits of agreement from −4.87 to +5.72 [123]. A study of AI scoring of endodontic radiographs found wide limits of agreement and systematic bias toward higher AI scores [157].
Classification Performance: Precision, Recall, F1, Sensitivity, Specificity
Classification metrics are used in two distinct contexts: binary or multi-class scoring tasks where the output is a category label, and pass/fail or flagging decisions. For automated detection of radiology report discrepancies, a fine-tuned LLM achieved F1 = 0.70, sensitivity = 66.3%, and specificity = 95.5%, outperforming other generative LLMs (Mixtral F1 = 0.60, Mistral F1 = 0.53, Llama2 F1 = 0.43) [9]. For binary classification of student physics responses as correct or incorrect, a conventional ML ensemble achieved F1 = 0.74–0.82 depending on the concept, while ChatGPT-based classifiers showed very low precision (0.16–0.20) [24]. For automated scoring of computer science feedback correctness, overall precision was 80%, recall 57%, and F1 67%, with inter-run variability noted [3].
In multi-class settings, macro-F1 is the preferred metric when class imbalance is present. A study of text complexity classification found macro-F1 of 45.44% for the best LLM configuration [183]. For classification of synchronization errors in student code, F1 values ranged from 0.09 to 0.59 depending on error type and model, with race condition detection performing best [33]. For classification of non-technical surgical skills as exemplar or non-exemplar, LLMs achieved macro-averaged F1 of 0.62–0.84 for exemplar behaviors but only 0.37–0.41 for non-exemplar behaviors, compared to classical ML F1 below 0.20 for the minority class [185]. For automated scoring of student prompts across binary dimensions, per-dimension accuracy ranged from 0.85 to 0.98, with an overall average of 0.92 [194].
A study of automated short-answer scoring for constructed-response items used a weighted RMSE (wRMSE) to address grade imbalance, finding GPT-4o achieved the best mean wRMSE of 0.27 across datasets [132]. For scoring of user stories against quality criteria, F1 scores ranged from 64% to 98.5% across criteria, with Conflict-Free and Unique criteria performing best [206].
Reliability and Consistency: Test-Retest and Inter-Run Stability
A recurring concern across the literature is the stochastic variability of LLM outputs when the same input is presented multiple times. Several studies explicitly quantify this. A study of essay grading found that ChatGPT produced scores of 91, 93, and 88 out of 100 across three iterations of the same essay, raising concerns about credibility [20]. A study of reflective essay and coding assessment scoring found coefficient of variation (CV) values averaging 13.57% for poor-quality reflective essays and 33.89% for poor-quality coding assessments, with consistency improving as the number of runs increased [2]. A study of physics solution grading found that inter-run variability differed substantially across models, with some models showing near-zero deviation across five runs and others showing larger variability attributed to temperature settings [61].
Several studies address this by averaging across multiple runs. A study of virtual patient interview scoring used five repeated runs and reported ICC(3,1) = 0.77 for single-run repeatability and ICC(3,5) = 0.94 for the average of five runs, compared to human inter-rater ICC(2,1) = 0.38 [123]. A study of history-taking dialog scoring found that after optimization, ICCs exceeded 0.923 across all scenarios, with CV values below 1.12% at the total-score level [162]. A study of essay scoring found ICC values of 0.999 for within-session consistency but a drop to 0.944 for a control measurement taken several months later, suggesting temporal drift in model behavior [48].
A study comparing 18 LLMs on programming assignment scoring found that ICC values against teacher grades ranged from 0.204 to 0.470, while most models achieved ICC above 0.8 against the model consensus, suggesting that models agree more with each other than with human teachers [117]. A study of medical SAQ grading found substantial within-model variability between two grading attempts, with kappa changes of up to ±0.50 for some model–question combinations [175]. A study of LLM scoring of writing found that DeepSeek showed a maximum within-essay score difference of 5 points across two runs, while Wenxin was more stable [165].
A particularly striking finding on evaluator-dependent bias comes from a study in which two AI evaluators (Claude and ChatGPT-4o) assessed the same writing outputs and produced completely reversed rankings, with Claude Sonnet 4 receiving 8 wins under its own evaluation but 26 wins under ChatGPT-4o’s evaluation—a 51.4 percentage-point swing—and complete agreement occurring in zero of seven domains [218].
Psychometric and Validity Evidence
Formal psychometric validation of AI grading systems is relatively rare in the literature but is present in a meaningful subset of studies. Classical test theory indices are the most common approach. A study of open-ended science item scoring found that Scenario 2 (rubric-guided) AI scores produced item discrimination values of 0.39–0.43 and item difficulty values of 0.52–0.67, closely matching expert values, while Scenario 1 (unguided) produced near-zero discriminations [34]. A study of automated short-answer scoring for a multilingual assessment found that item-total correlations from machine scoring (mean r = 0.33) were similar to those from human scoring (mean r = 0.35), supporting construct validity [21]. A study of AI scoring of student reflections used Rasch analysis to validate AI-generated scores, finding acceptable fit (residual factorization F1 = 1.74, below the threshold of 2.0) with one noisy item [203]. A study of creativity scoring using AI found that the LLM-based approach achieved Cronbach’s alpha of 0.92 across alternative exemplar sets and 0.98 across repeated evaluations with fixed anchors, indicating high internal consistency [182].
Predictive validity—the correlation of AI scores with external outcomes—is rarely reported. One exception is a study of a knowledge-graph-enhanced assessment system that found experimental group students showed greater learning gains (post-test means 78.4 vs 71.6, Cohen’s d = 0.56) and higher knowledge-mastery indices, providing evidence that AI assessments aligned with downstream learning outcomes [166]. A study of automated scoring of student reflections found that LLM-assessed scores improved downstream prediction of at-risk students compared to a text-only baseline [213].
Additional Metrics
Several studies report metrics that do not fit neatly into the above categories. Generalizability theory (G-theory) was used in one study to decompose variance in AI scoring, finding that 71% of total variance was attributable to differences among student answers and that the G-value for five AI sessions was 0.93, indicating high internal consistency despite poor agreement with human raters [150]. Krippendorff’s alpha is used in several studies as an alternative to kappa, particularly for ordinal data or multi-rater designs [doi:10.1016/j.mex.2023.102531; doi:10.26466/opusjsr.1821518; doi:10.1016/j.jacr.2025.12.024]. A study of translation quality assessment used BLEU scores to evaluate LLM output quality, finding improvement from 0.45 to 0.63 after fine-tuning [159]. A study of automated scoring of student prompts used Aiken’s V for content validity [167]. A study of singing quality assessment introduced a novel Human-in-the-loop Tiered Perceptual Ranking (H-TPR) metric, arguing that traditional accuracy metrics are invalid given annotator disagreement, and reported H-TPR values of 76.7%–82.9% for the proposed system [212].
Synthesis
Across the literature, QWK and Pearson/Spearman correlation are the most commonly reported metrics, reflecting the dominance of essay scoring and short-answer grading tasks where scores are ordinal or continuous. Cohen’s kappa and ICC are the preferred metrics in clinical and health professions contexts where rubric-based categorical agreement is the primary concern. MAE and RMSE are used most consistently in studies framing grading as a regression problem. Classification metrics (F1, precision, recall) are used primarily for binary or multi-class detection tasks rather than for continuous scoring.
The evidence base is strongest for structured, rubric-guided scoring of factual or procedural content, where multiple studies report QWK values above 0.80 and Pearson correlations above 0.85. The evidence base is weakest for tasks requiring interpretive judgment, higher-order reasoning, or evaluation of constructs such as reflection quality, clinical reasoning, and creative expression, where agreement with human raters is frequently poor and sometimes negative. Reliability across repeated runs is a persistent concern that is underaddressed in the literature: many studies report only a single run per item, and those that do assess inter-run consistency frequently find meaningful variability, particularly for lower-quality submissions. Formal psychometric validation—including IRT, Rasch, or predictive validity analyses—remains rare and represents a significant gap in the evidence base.
Outcomes
Learning benefits. A smaller but important set of augmentation studies documents downstream improvements in student outcomes attributable to faster, more consistent, or more personalised AI-generated feedback. The personalised programming assessment framework, which fine-tunes ChatGLM3-6B on expert-curated feedback instances and integrates a domain knowledge graph, reported a post-test mean of 78.4 versus 71.6 for controls (Cohen’s d = 0.56), with larger gains for initially lower-performing students and higher learner satisfaction (4.31 versus 3.21 out of 5) [166]. The GenAI grading explanation study observed modest transfer effects whereby educators previously exposed to AI insights showed some improvement in subsequent unaided assessment performance, though group differences in the transfer phase did not reach statistical significance [228].
Beyond the question of whether AI systems can match human judgment on scoring rubrics, a growing body of research examines what automated evaluation actually does to the broader educational ecosystem. This includes how AI tools affect the speed and cognitive burden of human evaluators, whether students who receive AI-generated feedback demonstrably learn more or differently than those assessed by humans, and — critically — whether the deployment of such systems can produce harms that aggregate accuracy metrics obscure. These harms range from systematic bias against particular student populations to the erosion of pedagogically valuable feedback practices, and they represent an important counterweight to optimistic claims grounded solely in inter-rater reliability statistics.
Evaluator Efficiency
The efficiency gains documented across the literature vary considerably in magnitude, precision, and domain, but collectively paint a picture of AI evaluation systems that can dramatically reduce the time and human effort required to assess student work — while also revealing important tensions between speed and accuracy.
The most precisely quantified time savings appear in studies of clinical and medical education contexts. In one OSCE grading study, a customized AI system graded all 39 students at one station and produced individualized feedback in under five minutes, compared with approximately eight hours of faculty time for the same station during a live examination; projected across five stations, AI grading was estimated to require under ten minutes total [115]. A comparable finding emerged from a study of virtual patient interview transcript scoring, where AI systems reduced per-case processing time by 58–68% relative to physician raters, translating to throughput increases from approximately six to fourteen or eighteen cases per hour and freeing an estimated 210–240 faculty minutes across a 35-case set [123]. In pharmacy education, automated grading of SOAP notes required only nine minutes total for 83 notes (approximately 0.1 minutes per note) compared with 415 minutes for faculty graders (five minutes per note) and 4,150 minutes for peer graders with written feedback — a reduction exceeding 90% in raw grading time, though the authors noted that AI setup and faculty oversight time were excluded from these figures [158].
Comparable efficiency gains have been documented in writing and essay assessment. One study reported that ChatGPT graded and commented on an essay in approximately one minute versus approximately 45 minutes for a human lecturer, characterizing this as 45 times faster than human capacity and projecting that 100 essays would require roughly 100 minutes of AI time versus 75 hours of lecturer time [36]. A large-scale writing assessment study processed 153 writing samples in approximately 1.5 hours using ChatGPT-4, compared with an estimated several dozen hours or a traditional timeline of two to four weeks for human-only assessment, with the AI integration shifting human effort from labor-intensive scoring toward prompt design and oversight [187]. In a chemistry education context, AI-assisted grading was estimated to absorb between approximately 47% and 85% of the grading workload depending on the confidence measure used, with the authors framing this reduction as enabling a return to more open-ended assessment formats by lowering the teaching assistant days required for grading [188].
Automated systems for short-answer and open-ended question grading have similarly demonstrated substantial throughput advantages. A retrieval-augmented grading system for open-ended responses processed a 90-exam validation set in approximately 40 minutes of computation (roughly 0.5 minutes per exam on a consumer GPU) compared with approximately 25 hours of manual instructor time, with the authors projecting that overall instructor time could be reduced from 25 hours to approximately two to three hours when including review of flagged cases [138]. A multi-agent feedback framework reported an average feedback generation time of 12.3 seconds per item versus an expert manual average of approximately 1,847 seconds, yielding a roughly 99.3% reduction in generation time while retaining approximately 92% of expert-level feedback quality; the study also provided cost estimates of approximately $1.80 per student per semester for full deployment [166]. For programming assignment grading, one study reported that automated grading via the OpenAI API took approximately 15–20 seconds per assignment, compared with an estimated five to ten minutes per submission for human graders, with total automated grading of 534 assignments requiring approximately 2.5 hours versus an estimated 45 or more hours of human effort — a reduction of over 90% — at an average cost under $0.05 USD per graded assignment [169].
Automated scoring of reflective writing has also been timed precisely. A multi-agent pipeline for rubric-based scoring of student reflections averaged 7.71 seconds per reflection for scoring alone and approximately 33.35 seconds for the full feedback pipeline, compared with a human grader average of 1.4 minutes per reflection — approximately eleven times faster than the human average and twenty times faster than the slowest human evaluator. The per-reflection monetary cost was estimated at approximately $0.0015, with 84 reflections costing approximately $0.13 in total [211]. In a Japanese language writing context, instructors reported that AI-assisted OCR and automated feedback significantly reduced marking time and enabled rapid generation of error summaries, though no quantitative time figures were provided [170].
Several studies report efficiency gains in terms of reduced video review burden rather than grading time per se. A video segmentation system for OSCE physical examination identification reduced the mean video requiring human review from 900 seconds to 175 seconds per encounter — an 81% reduction in video needing review — by automatically identifying and extracting relevant examination segments [204]. A clinical simulation platform handling 1,024 encounters across 410 learners demonstrated practical scalability, with automated rubric-aligned assessments typically completing within two to three minutes post-encounter, enabling rapid formative feedback and relieving faculty of first-pass scoring [199].
Cost comparisons between AI and human annotation have been reported in a small number of studies. One study found that fine-tuning GPT-3.5 for narrative feedback quality scoring cost approximately $76.83 in total API and fine-tuning expenses, compared with human annotation costs ranging from $1,000 to $15,000 (with the human-labeled dataset for that study costing approximately $3,700) [196]. Another study reported that LLM grading cost approximately $0.10 per student per homework assignment for Claude and substantially less for Gemini, while also noting that automated video analysis for academic integrity flagging reduced human verification time from approximately one hour to approximately one minute of targeted inspection per recording [215].
Many studies report efficiency gains only qualitatively or through inference rather than direct measurement. Authors frequently assert that AI grading can reduce teacher workload, enable scalable formative assessment, or provide faster feedback, without providing timed comparisons or cost estimates [54], [55], [62], [53], [49], [129], [116], [161], [163], [185], [144]. In language learning contexts, teachers reported an average of 18 minutes saved per class through AI-assisted pronunciation feedback, with additional time reductions in assessment integration and parent-teacher reporting, though these figures were self-reported rather than objectively measured [143]. Similarly, a study of AI-assisted short-answer grading reported that 88% of questions were graded with AI assistance, enabling personalized feedback at an average class size of approximately 70 — something instructors would otherwise provide only with a class a third or a quarter as large — though no precise time-savings figures were given [120].
An important tension between speed and accuracy runs through several of these efficiency reports. Studies that document the fastest automated processing times do not always report the highest accuracy, and some explicitly acknowledge that efficiency gains depend on the quality of the underlying system. One study found that smaller language models were nearly instantaneous (under 0.02 seconds per item) compared with approximately 7.5 seconds for GPT-based systems, but at a substantial cost in scoring accuracy — a direct speed-accuracy trade-off [153]. Another study found that natural-language grading insights improved grading correctness and feedback quality but did not meaningfully speed up grading and in some cases required more time than simpler approaches, suggesting that qualitative improvement and efficiency do not always move together [228]. A study of automated scoring of NAEP mathematics items noted that the approach could enable scoring “cheaply and at scale” but did not provide measured cost or time data to substantiate this claim [14]. Similarly, a study of automated scoring of programming assignments reported that AutoSCORE yielded better scoring performance but at the cost of increased inference time per instance, explicitly framing this as an efficiency-performance trade-off [140].
Scalability to larger cohorts is frequently cited as a motivation for AI evaluation systems, particularly in contexts where human grading is resource-constrained. Automated scoring of bioinformatics assignments handled approximately 105 submissions per assignment with a manual review request rate of only 0.6%, suggesting low human overhead for post-hoc review at scale [136]. A radiology discrepancy detection system demonstrated that LLM curation could enrich the prevalence of high-teaching-value cases from 10.4% to 85.8%, substantially improving the efficiency of case selection for educational use even without direct measurement of time saved [9]. A creativity assessment platform enabled one-click autoscoring of collected data and CSV-based batch scoring without programming expertise, supporting scalable deployment across multiple creativity assessment tasks [148].
Across the literature, the most credible efficiency claims are those accompanied by direct timing comparisons, cost estimates, or throughput measurements. Studies that rely solely on qualitative assertions of efficiency gains — while numerous — provide weaker evidence for the magnitude of actual workload reduction. Where both efficiency and accuracy data are available, the literature suggests that the realized efficiency benefit depends critically on the accuracy of the underlying system: a fast but inaccurate grader may generate additional human review burden that offsets the initial time savings.
Learning and Other Outcomes
The evidence on student learning effects, perceptions, implementation challenges, and other outcomes is extensive and frequently mixed, reflecting the diversity of contexts in which AI evaluation has been deployed.
Student Learning Effects
Several studies report measurable gains in student performance associated with AI-generated feedback. A quasi-experimental study of nursing students found that those with access to an LLM feedback system achieved significantly higher pitch scores than controls (mean 19.68 vs. 17.30, Cohen’s d = 0.73) [161]. A knowledge-graph-augmented LLM system similarly produced higher post-test scores (mean 78.4 vs. 71.6, Cohen’s d = 0.56), with particularly pronounced gains for initially lower-performing students and improved engagement indices [166]. Students using AI-supported feedback in a writing context achieved a higher mean final score (87.8% vs. 81.2%), with bottom-quartile students gaining approximately ten percentage points, and reported increased confidence and engagement [128]. An AI-assisted reflective feedback system was associated with medium average normalized gain and large effect sizes in student conceptual understanding [167]. In a clinical simulation context, one institution showed a significant improvement in OSCE history-taking scores (mean 82.8 to 88.8, d = 0.75) following use of an AI-based oral examination tool, though a parallel pilot at a second institution showed no significant difference [199]. A study of LLM-assisted formative testing found higher mean scores for the LLM-assisted group on both formative (64.16 vs. 55.6) and summative assessments (27.8 vs. 26.08) [53]. An automated pronunciation assessment system produced larger pronunciation gains and strong improvements in self-regulated learning and metacognitive knowledge (Cohen’s d up to 2.14) [143]. Students using a creativity scoring platform showed increases in fluency and flexibility with practice, though effects on originality and elaboration were more complex and limited [10].
By contrast, other studies found no significant learning advantage. A study comparing GPT-4 feedback to human tutor feedback in L2 writing found both groups improved over time but no significant time-by-group interaction (F = 3.094, p = 0.085) and no between-group effect [4]. A study of student prompting literacy found significant improvement only in the Background dimension of prompt construction, with no significant changes in other dimensions [194].
Instructor and Student Perceptions, Trust, and Satisfaction
Perceptions of AI evaluation are consistently mixed and context-dependent. The most detailed experimental evidence on trust comes from a study in which students were asked to evaluate ChatGPT-only, instructor-only, and hybrid grading: trust was highest for instructor grading (M = 6.29), lowest for ChatGPT alone (M = 4.29), and intermediate for the hybrid condition (M = 5.50); no student preferred ChatGPT-only grading, and 15 of 24 preferred the hybrid [56]. Ethical and benevolent dimensions of trustworthiness increased significantly after the experience, but performance-related dimensions (reliability, competence, transparency) did not change significantly, and intent to rely on ChatGPT remained modest [56]. A matched pre/post-reveal survey of 158 students showed that satisfaction, trust, and perceived fairness all increased after students learned ChatGPT had been used for grading, with a strong effect size (r = 0.913), though some students still preferred human grading for perceived leniency or contextual sensitivity [169].
In dental education, students reported no significant differences between human and AI feedback on clarity, relevance, usefulness, promotion of critical thinking, or perceived improvement of future performance, though they reported greater comfort with human feedback (χ² = 9.01, p < .05); experts rated AI feedback higher on identifying mistakes and suggestions for improvement [25]. Students in a clinical simulation study rated expert feedback higher on average but differences were not statistically significant, and 53.1% preferred a combined expert-plus-AI approach [157]. Radiology residents showed moderate satisfaction with LLM-generated feedback (mean 3.50/5) and a majority (71.43%) preferred a hybrid approach [31]. Students using an automated medical training evaluation system reported strongly positive attitudes: 87% found the system helpful, 83% wished to use it in the future, and 90% would recommend it [162]. Students using a chatbot-based formative assessment tool reported high satisfaction (89.17%) and strong willingness to recommend it (82.5%) [65]. A nephrology OSCE preparation chatbot achieved high usability scores (91.7% rated it ≥7/10 for preparation) and strong ratings for feedback helpfulness and personalization [163]. Students in a bioinformatics course showed no significant overall preference between LLM and human TA feedback, though one model produced worse ratings for incorrect answers [136]. In a Russian-language pilot, only overall satisfaction significantly favored instructors over AI (Cohen’s d = 0.55), with no significant differences on other quality criteria [203].
Perceptions were more polarized in other contexts. A study of AI evaluation in a Russian educational platform found approximately balanced positive, negative, and neutral sentiment, with positive perceptions centered on rapid availability and personalization and negative perceptions centered on inconsistency, perceived formalism, occasional excessive strictness, and erosion of trust [122]. Students in a qualitative study reported difficulty understanding how final grades were determined when AI was involved, reducing perceived transparency and institutional trust, and expressed concerns about fairness for non-standard linguistic styles [209]. A focus-group study of faculty and students identified four themes: trust with skepticism, fairness concerns (perceived bias toward polished English), irreplaceability of teacher judgment, and expectations for transparency and governance [146]. A survey of educators and students found mixed opinions about using GenAI to mark assessments (moderate skepticism overall), greater comfort when AI feedback was combined with instructor feedback, and highly variable experiences with AI detection tools including reported false positives with consequential impacts [193].
Instructor perceptions were similarly varied. Teachers in a programming assessment study judged ChatGPT grading as not coherent (mean 1.68/5), were unwilling to change their own assessments based on it (1.42/5), and did not consider ChatGPT-only assessment reasonable (1.89/5) [46]. By contrast, a lecturer in an EFL context reported full agreement with ChatGPT’s scores, was impressed by its comments, and expressed willingness to use it in assessment, while emphasizing it cannot fully replace teachers [36]. Teachers in a study of AI-assisted pronunciation assessment reported enhanced diagnostic capability (23/24) and improved classroom efficiency (22/24), with high satisfaction ratings [143]. An institutional assessment committee formally decided to adopt GenAI for future written communication assessments after reviewing outputs [187]. Educators rated AI-generated natural-language grading insights significantly higher than important-word highlights on informativeness, comprehensibility, and willingness to adopt [228]. Teachers in a study of AI writing evaluation were generally cautiously positive, seeing strengths in surface-level features but weaknesses in nuanced judgment, and emphasizing the need for human oversight [171]. A study of music performance grading found mixed expert responses, with some teachers praising ChatGPT for rapid rubric generation and others flagging music-specific inaccuracies [181]. A survey found 71.79% of respondents supported autonomous AI assessment and 54.70% supported allowing AI in assessments, though awareness of institutional AI policies was low (47.86%) [134].
Students’ perceived helpfulness of AI feedback was also examined behaviorally: in a large-scale study, 60.5% of students pooled across conditions rated feedback as helpful, with human grading modestly increasing perceived helpfulness overall and more strongly for lower-ACT/SAT students [120]. Students in a clinical training context rated AI-generated feedback as most helpful in 62.5% of instances when models were trained on domain-specific data, citing clarity and alignment with clinical reasoning, while untrained models were rated as most helpful in only 9.4% of instances [216]. Students in a programming learning platform reported high perceived fairness of AI grading (87.4% rated it 4–5/5) and willingness to continue using the system rose from 74.1% to 87.5% [223].
Implementation Challenges
Practical obstacles in deploying AI evaluation were reported across many studies. Prompt sensitivity and rubric design emerged as pervasive challenges: grading quality varied substantially with prompt formulation, rubric specificity, and whether rubrics were provided at all [2], [54], [34], [37], [175], [181]. Providing rubrics did not consistently improve performance and sometimes degraded it [175]. Rubric sensitivity was also demonstrated by showing that grading the same solution under different rubrics yielded different scores [169].
Hallucinations and mathematical errors were identified as significant implementation risks. LLMs exhibited hallucinated errors, presented submitted code as corrected solutions, and ignored assignment instructions [3]. Mathematical hallucinations, failure to recognize algebraically equivalent solutions, and over-leniency were documented in physics grading [61]. Arithmetic errors and loss of coherence in extended tasks were noted in mathematics grading [43]. OCR inaccuracies affected handwriting-based evaluation [170].
Context limitations were widely reported: inability to supply full project context led to generic suggestions and reduced capacity to detect deep logical or structural issues in software engineering [5]. AI models evaluated only textual transcripts and could not assess nonverbal, auditory, or visual cues, producing systematic divergence from human raters in clinical and performance contexts [121], [222], [177], [141]. Modality limitations also affected circuit diagram interpretation [142] and visual design evaluation [154].
Institutional and governance challenges were prominent. Uneven preparedness across institutions, fragmented governance, limited training, and reliance on individual lecturers led to inconsistent practices [209]. Data governance and privacy concerns were widespread, with uncertainty about how student submissions processed by LLMs are stored or reused [209]. Aligning AI tools across diverse curricula required significant time investment for staff training and continuous updates [60]. Technical issues including system latency, login problems, and voice-interface inaccuracies were reported in several deployments [194], [224], [199]. Performance degradation beyond 70 concurrent users was noted in one system [100]. ChatGPT Plus experienced a three-hour lockout after approximately 80 file uploads in one study [158]. Session and data limits requiring session restarts were noted in music grading [181].
Run-to-run variability and non-determinism posed reliability challenges across many contexts [2], [3], [48], [173]. Systematic score differences between AI and human evaluators—with AI tending toward either leniency or strictness depending on context—were documented in dental [221], medical [47], [178], clinical communication [121], and writing assessment [155], [135] contexts. Systematic evaluator effects producing complete ranking reversals across tools were identified as a threat to assessment equity [218].
Other Reported Outcomes
Several studies reported outcomes that do not fit neatly into the categories above. A study using NLP to analyze student writing over time identified a progressive decline in rhetorical and linguistic quality that the authors mapped to markers of deteriorating mental health, proposing a trauma-informed pedagogy framework and early-warning system [17]. An LLM-curated case set for radiology resident education contained a higher prevalence of discrepant and educationally valuable cases than a random set, suggesting potential utility for trainee oversight [9]. A study found that 7.7% of MCQs produced completely false responses from both AI models, prompting faculty review and identification of ambiguous or poorly constructed questions, suggesting a secondary benefit of AI evaluation for quality assurance of assessment instruments [156]. Students in a programming course who used AI tools reported decreased dependence on LLMs over the semester, improved prompt-crafting, and improved verification strategies, with documentation requirements positively impacting learning [215]. A study of student code submissions found that approximately 12% were flagged as anomalous by a behavioral clustering approach, with student reflective notes showing generally positive attitudes toward AI guidance alongside concerns about fairness and critical thinking [229]. An AI-assisted code review workflow stimulated greater student participation (70% of reviews used GenAI), prompted students to research and implement security fixes, and led some students to continue using ChatGPT beyond course requirements, though some found feedback generic [5]. Students’ self-reported confidence in using AI for learning increased by 10.4% after an AI-based prompting literacy activity [194]. One study reported that some students over-relied on AI feedback, potentially limiting independent learning [128], while another noted that students sometimes trusted ChatGPT’s interpretations of code even when those interpretations were questionable [5]. A study of AI-assisted teacher assessment found that teachers assigned significantly higher grades in AI-mediated viva sessions than in face-to-face sessions, though with a small effect size [224]. Concerns about widespread ChatGPT availability undermining the validity of assignment-based assessment were raised in the context of programming education [46].
Inversion: Cases Where AI Evaluation Introduced Harm
The most pervasive inversion signal across the literature is poor agreement with human graders, a pattern that cuts across subject domains, educational levels, and AI architectures. Several studies document agreement so low as to raise fundamental validity concerns. A study deploying ChatGPT-4o to grade short-answer responses in a postgraduate mechanical ventilation course found that Cohen’s kappa between AI and human raters was −0.079, that 63% of student answers showed unacceptably large discrepancies, and that the AI systematically scored lower than the human grader by a mean of 1.34 points on a 10-point scale — a bias toward false negatives that the authors concluded rendered the system unsuitable for high-stakes assessment [150]. A parallel finding emerged from a study applying ChatGPT to reflective essay scoring, where Cohen’s kappa was −0.048 and ICC values were negative, accompanied by a systematic positive bias in which the AI consistently awarded higher scores than human raters [155]. A study evaluating ChatGPT-4o for OSCE performance assessment found ICCs close to zero or negative across all six evaluated domains, with the AI assigning significantly inflated scores relative to physicians in four of six domains [222]. Similarly, a study grading medical students’ SOAP notes found that ChatGPT-4 awarded honors grades to 92.9% of students compared to 63.8% awarded by the human proctor, with restricted cubic spline modeling explaining less than 10% of variance in human scores [178]. A study using ChatGPT-4o to evaluate root canal treatment radiographs reported Cohen’s kappa of 0.210 and Pearson correlations below 0.30 across all five assessed criteria [221], while a companion study using ChatGPT-4o for the same task found ICCs ranging only from 0.36 to 0.45 [157].
Agreement failures are not confined to medical education. A study deploying ChatGPT-3.5 as an automated essay scorer found ICC values below 0.50 for most rubric criteria, with the AI consistently awarding substantially higher scores than the human rater across all dimensions — a large-effect-size leniency bias (η² up to 0.80 for grammar and spelling) that the authors warned would impair discrimination between low- and high-performing students [135]. A study evaluating ChatGPT-3 on a self-generated essay found that the human researchers scored the same essay 41/100 while ChatGPT awarded itself 88–93/100 across three runs, with the rubric itself containing arithmetic errors and the scores varying across repeated evaluations [20]. A study grading calculus exercises found that while GPT-4’s average scores aligned with human graders at the aggregate level, per-item aptness indices and correlations across runs (ranging from 0.19 to 0.69) revealed persistent individual-level unreliability, with the system making frequent elementary arithmetic errors and suffering loss of coherence over long grading sequences [43]. A benchmarking study of 18 LLMs grading programming assignments found ICC values against teacher grades ranging only from 0.204 to 0.470, with the authors explicitly noting that agreement with human teachers remained low-to-moderate across all tested models [117]. A study evaluating LLMs for detecting synchronization errors in student Java code reported overall accuracy of approximately 50%, with precision values as low as 0.05 for some error types and configurations [33]. A study applying ChatGPT-3.5 and Google Bard to rate prompt complexity found ICCs of 0.22 and 0.05 respectively against human raters, compared to a human-only ICC of 0.84 [19]. A study evaluating six general-purpose LLMs on language exam grading found Cohen’s kappa values ranging from −0.164 to 0.344 against a human gold standard, with RMSE values between 4.69 and 8.49 on a 100-point scale compared to 1.41 for human graders [176].
In clinical and health professions education, inversion patterns are particularly pronounced. A study deploying ChatGPT-4o and Gemini Flash 1.5 to assess OSCE procedural skills via video found that inter-rater reliability fell below acceptable thresholds across many criteria, with AI models consistently assigning higher scores than human raters and showing larger systematic biases on auditory and nuanced tasks [47]. A study applying ChatGPT-4o to prehospital simulation transcripts found that weighted Gwet’s AC2 remained below the pre-specified threshold of 0.60 across all five non-technical skills domains, with two domains yielding negative AC2 values, and that the AI tended to award higher scores than faculty in three domains — a pattern the authors attributed to the system’s emphasis on surface linguistic features over grounded clinical appropriateness [121]. A study using GPT-4 to evaluate radiology report discrepancies achieved sensitivity of 79.2% with moderate inter-rater reliability (Fleiss’ kappa = 0.43) among resident evaluators, and the system occasionally flagged clinically insignificant findings as missed diagnoses [31]. A study using ChatGPT-4o to assess surgical video quality via transcripts found that AI runs were internally consistent but assigned significantly lower scores than human evaluators across both DISCERN and GQS instruments, a systematic underestimation the authors attributed to the model’s inability to process audiovisual features [177]. A study applying VideoLLMs to science-popularization video assessment found ICC values below 0.40 for most evaluated measures, with one model systematically overestimating quality and another systematically underestimating it [141].
A second major inversion pattern involves construct-irrelevant scoring — cases where AI systems reward or penalize features that are not the intended targets of assessment. A study grading physics problem solutions found that LLMs in blind mode showed substantial leniency, awarding high marks to incorrect or incomplete solutions and sometimes failing to recognize equivalent correct expressions, with mathematical hallucinations identified as a principal cause [61]. A study evaluating GPT-4 and Gemini for short-answer grading in medical education found that both models tended to over-score incorrect answers, with ChatGPT-4 showing increased errors on verbose or disorganized responses — a pattern consistent with length-based heuristics — while also exhibiting high rates of pass-fail misclassification near cut-scores (ChatGPT-4 flip rate 51.5%) [144]. A study grading dental histology responses found that the AI sometimes awarded credit for content beyond the rubric, including examples not specified in the marking criteria, and aligned with the human tutor only approximately 50% of the time [25]. A study applying GPT-4 to automated essay scoring in oral and maxillofacial surgery found that the system sometimes awarded marks for factually incorrect or irrelevant content while simultaneously producing lower mean scores than human graders on one of two questions, suggesting simultaneous false-positive and false-negative tendencies [54]. A study evaluating ChatGPT-4o for scoring handwritten open-ended test responses found that without rubric guidance, Cohen’s kappa was approximately 0.09 and correlation with human scores was 0.12, with the AI awarding full credit to inflated or off-topic responses and applying excessive strictness to brief but correct answers [34]. A study using Claude 3.5 Sonnet to evaluate student summaries found near-zero and non-significant correlations for the paraphrasing criterion (r = 0.12), with the model failing to distinguish lexical copying from structural paraphrasing and not penalizing unchanged technical terms [147]. A study applying ChatGPT-3.5 to classify physics concept use in student sentences found that the LLM-based classifier achieved precision values of only 0.20 and 0.16 for the two target concepts, with the model over-assigning positive labels and sometimes scoring based on the mere presence of concept-related strings rather than genuine conceptual application [24].
A third cluster involves inversion through unreliability across repeated runs, which undermines the basic requirement that assessment be consistent. A study using ChatGPT-3.5 to grade programming assignments found that identical submissions received scores as different as 10 and 0 across repeated runs, leading the authors to conclude that this inconsistency alone disqualified the system as a standalone assessor [46]. A study evaluating a GPT-4-based grading system for calculus exercises documented temporal fluctuations and low inter-run correlations, with the system additionally suffering from hallucinations and loss of coherence over long grading sequences [43]. A study using a customized ChatGPT for music essay scoring found that three repeated runs produced total score means of 10.52, 9.88, and 11.40 for the same submissions, with a maximum within-essay score difference of five points observed for DeepSeek across runs [181]. A study deploying GPT-3.5 to evaluate argumentative writing found variability in precision and low recall for recognizing complete works, with the system exhibiting a conservative classification tendency that systematically underestimated article quality [6]. A study using a GPT-4-powered marking system for coding and reflective essay tasks found coefficient of variation values as high as 33.89% for poor-quality coding submissions and documented cases where the system awarded high marks to flawed work while generating justifications that contradicted the assigned scores [2]. A study using GPT-3.5 to classify student code for synchronization errors found that run-to-run variability was present, with accuracy on some tasks dropping to 57% in individual runs [3]. A study using an LLM-based Moodle grading tool found that users reported the same answer being graded differently across attempts, with negative reviews focusing on grading inconsistency as the primary complaint [122].
A fourth inversion pattern involves cases where AI systems fail to capture higher-order or domain-specific constructs, effectively measuring something other than the intended competency. A study applying GPT-4 to score Gibbs-structured reflective writing found that while the system showed strong alignment with human raters on surface-level dimensions such as clarity and language mechanics (ICC ≈ 0.83–0.86), agreement dropped substantially for higher-order dimensions including feelings (ICC = 0.52) and analysis (ICC = 0.59), with participants reporting that the AI privileged polished English over substantive reflective depth [146]. A study evaluating ChatGPT-3.5 for essay classification found that humans were approximately 2.57 times more likely to correctly classify texts than the AI, with the system efficient at identifying surface-level errors but struggling with contextual subtleties and stylistic complexity [149]. A study applying ChatGPT to evaluate student writing on an EOP platform found that while the system performed adequately on grammar detection, it missed context-specific meaning and creative or complex ideas, with teachers reporting that it favored shallow organization over deep logical thinking [171]. A study using a GPT-4-family model to score reflective writing found that AI–human correlations for higher-order dimensions such as action plan (r = 0.59) and analysis (r = 0.54) were substantially lower than for surface dimensions, with the overall AI–human ICC of 0.71 falling below the human inter-rater ICC of 0.82 [146]. A study using two domestic LLMs to grade high school English compositions found that while Wenxin achieved moderate correlations for numerical scores (r = 0.748), DeepSeek’s grade-level correlation was only 0.236, and both models were better at identifying surface errors than assessing topic adherence or emotional expression [165].
Several studies document inversion through false positives that could directly harm evaluation integrity. A study using a task-customized GPT to assess A1-level Spanish writing found that the system frequently generated false negatives — flagging correct sentences as erroneous — and produced vague or irrelevant feedback, with none of the eight tested versions meeting human reliability benchmarks and identical versions producing different scores across runs [173]. A study evaluating five LLMs for medical short-answer grading found highly variable kappa values across questions and models, with a general tendency toward conservative grading producing higher false negative rates, and with providing expert rubrics sometimes worsening rather than improving performance [175]. A study using GPT-4 and Gemini for short-answer grading found that all three tested LLMs showed reluctance to award zero points to clearly incorrect answers, producing a systematic false-negative pattern in which LLMs deducted points where humans gave full credit [49]. A study using a RAG-based automated assessment system for concept-map scoring found quadratic weighted kappa values near zero or negative for small-LLM configurations, with the best-performing configuration (FAISS-GPT) achieving only a macro-F1 of 0.338 [153].
A fifth inversion type involves automation bias — cases where the deployment context creates conditions for human evaluators to defer inappropriately to AI outputs. A study examining AI text detectors alongside generative AI feedback tools documented multiple instances where assessors lowered student grades based on detector outputs without adequately reviewing the student work, and where students were falsely accused of AI use by detectors that produced misleading results [193]. A study using GPT-4 to evaluate argumentative writing noted that the AI’s apparent speed and confidence could encourage over-reliance, with the authors cautioning that black-box AI judgments might be accepted without adequate scrutiny [20]. A study using an LLM-based Moodle grading tool found that users raised concerns about epistemic trust and inappropriate delegation, with the authors recommending hybrid models and teacher involvement to prevent blind reliance on automated outputs [122].
Several additional studies contribute evidence of inversion in specialized contexts. A study applying transformer-based models to classify narrative feedback into ACGME competency categories found that no transformer model outperformed the FastText baseline, with BERT-tiny performing worse than the baseline [41]. A study using an LLM to analyze student experimentation protocols found accuracies ranging from 0.38 to 1.00 across error categories, with the system struggling particularly with variable nomenclature and synonym recognition [11]. A study using a multimodal system to score classroom observation components found that GPT-3.5 achieved a near-zero correlation with human ratings (r = 0.027), though an ensemble combining GPT-4 with supervised models approached human inter-rater reliability [12]. A study using a GenAI pipeline to score creativity on the Alternative Uses Test found moderate correlations with human scores (r ≈ 0.57–0.59), with the authors noting potential score inflation and distributional biases in LLM-based originality scoring [10]. A study using a BERT-based self-augmentation method for automated short-answer grading acknowledged that augmented sentences may not be semantically similar to originals, risking improperly trained models on the small educational datasets typical of this domain [45]. A study evaluating a hybrid deep learning model for generating marking schemes found token-level accuracy of only 0.2632 and macro-F1 of 0.1818, though high BERT similarity scores (0.9944) suggested semantic alignment despite lexical mismatch [195]. A study using an AI-based grading pipeline for translation quality assessment acknowledged that machine translation cannot fully replace human expertise and that generative models may omit key information, leading to incorrect negative judgments [159]. A study using an NLP-based system to detect trauma indicators in student writing raised concerns about population bias, limited reproducibility, and the risk that surface-level linguistic features could be mistaken for indicators of genuine distress [17]. A study using an LLM-based short-answer grading benchmark found that even the best-performing system (GPT-4o, wRMSE = 0.27) had error more than double that of human graders [132]. A study using a GPT-4-based system to extract clinical experiences from medical student learning logs found overall sensitivity of only 62.39%, with category-specific sensitivity as low as 45.43% for symptoms — a false-negative rate that could systematically undercount students’ documented clinical exposure [124]. A study using seven AI tools to evaluate student website projects found weak or inconsistent Spearman correlations between AI and human instructor ratings for many criteria, with some tools systematically over- or under-scoring specific dimensions [154]. A study using ChatGPT to evaluate dissertation questions found score discrepancies in both directions relative to human graders, with the system awarding high scores to responses labeled partially correct by prior evaluators [180]. A study using a singing assessment system found inter-expert exact agreement below 45% across all dimensions, motivating a shift away from score-matching metrics entirely [212]. A study using AI raters to evaluate C2-level Turkish translations found an overall Krippendorff’s alpha of 0.392, with one model rewarding summaries that violated task instructions and the other applying systematically harsher penalties than the human expert [208]. A study using ChatGPT to generate marking memoranda found that the system failed on computational questions in engineering, producing a build-time estimate of 2.5 hours against a correct answer of 7.82 hours, while AI-generated student text received a 0% similarity score from Turnitin — illustrating simultaneous failures in grading accuracy and academic integrity detection [219]. A study using LLMs to classify Portuguese text complexity found accuracy and macro-F1 values often below 50%, with models performing worse than baselines in some zero-shot configurations and showing high run-to-run variability with random example selection [183]. A study using a GPT-4-based system to evaluate SQL query solutions found that models sometimes rewrote student responses by inserting clauses not present in the original submission, and that Copilot failed to grade three questions entirely in the initial approach [37]. A study using ChatGPT to perform automated code review found that the system identified significantly more issues than human peers (mean 26.7 vs 9.5, p = 0.01) but missed more critical structural and logic issues, with a large standard deviation indicating high inconsistency across teams [5]. A study using a circuit-analysis homework grading benchmark found that GPT-3.5 achieved only 51.66% average correct rate across metrics, with all models failing to recognize equivalent notations and GPT-4o hallucinating non-existent circuit components [142].
Evaluated Risks and Remediation
As generative AI systems take on increasingly consequential roles in educational assessment, the formal evaluation of their associated risks has emerged as a critical area of scholarly inquiry. Ensuring that automated grading is not only accurate but also fair, transparent, and robust requires systematic scrutiny that goes well beyond simple performance benchmarking — encompassing concerns such as demographic bias, susceptibility to manipulation, and the downstream effects of errors on student outcomes. Accordingly, the literature in this space spans a wide methodological range, from retrospective statistical analyses that detect disparate impacts across student subgroups to proactive adversarial testing regimes designed to expose vulnerabilities before systems are deployed in high-stakes contexts.
Types of Risks Evaluated
The most consistently evaluated risk across the literature is low agreement or accuracy relative to human grading. A substantial majority of studies formally measured AI-human agreement and found it to be present as a concern, though the severity varied considerably by task type, model, and prompting strategy. Studies using essay scoring, short-answer grading, and clinical assessment frequently reported moderate or poor agreement. For instance, ChatGPT-based graders showed poor inter-rater reliability with human raters when scoring IELTS prompt complexity [19], reflective writing [146], and prehospital simulation non-technical skills [121], while a study of postgraduate short-answer grading found ICC values near zero and over 60% of answers falling outside acceptable discrepancy bounds [150]. In medical education contexts, ChatGPT-4 showed significantly different scores from physicians across most OSCE domains [222], and agreement between AI and human graders for dental radiographic assessment was poor to moderate [157]. Physics problem grading showed substantial leniency and low agreement in blind conditions across multiple LLMs [61], and calculus grading with GPT-3.5 and GPT-4 produced frequent arithmetic errors and hallucinations [43]. Agreement was also found to be task-dependent: some studies reported acceptable agreement for structured or objective items while finding poor agreement for subjective or higher-order criteria [49], [163], [47]. A number of studies found agreement to be acceptable or high after careful prompt engineering, rubric provision, or fine-tuning [14], [129], [138], [139], [162], while others found that even strong overall correlations masked individual-level unreliability [46]. Across the full body of evidence, low agreement is confirmed as a real and widespread risk, though one that is partially amenable to mitigation through rubric alignment, fine-tuning, and ensemble approaches.
False positives — cases where AI graders awarded undeserved credit to incorrect or low-quality responses — were formally evaluated in a large number of studies and confirmed as present in most of them. The pattern of AI leniency or generosity was documented across diverse contexts. ChatGPT awarded high scores to a mediocre essay that human raters scored much lower [20], and consistently assigned higher scores than human raters in oral assessment of OSCE skills [47], automated essay scoring [135], and SOAP note grading [178]. In medical short-answer grading, AI models tended to over-score incorrect answers, producing critical pass/fail misclassifications [144]. In the context of open-ended physics and coding tasks, AI graders sometimes awarded marks despite missing key components or failing to detect errors [61], [2]. In multilingual short-answer scoring, machines sometimes scored responses as correct when humans scored them incorrect [21]. A study of ChatGPT grading reflective essays found a systematic positive bias (mean difference +0.497) and statistically significantly higher scores than human raters [155]. In automated essay scoring of nursing pitch scripts, models generated feedback that inaccurately claimed elements were present when they were missing [161]. False positives were also documented in code review [5], programming assignment grading [3], and translation quality assessment [159]. A smaller number of studies found false positive rates to be low or mitigated, particularly when high specificity was achieved [124], [192], or when confidence gating and calibration were applied [188].
False negatives — cases where AI graders incorrectly penalised correct or high-quality responses — were also widely confirmed. Several studies found AI systems to be systematically stricter than human graders. A study of medical short-answer questions found a higher rate of false negatives across all tested LLMs, with a conservative bias that may unfairly penalise students [175]. In postgraduate short-answer grading, AI was systematically harsher, with a mean difference of −1.34 points and 63% of answers showing unacceptably large discrepancies [150]. Programming assignment graders showed low recall for some error types [33], and LLMs grading student debugging problems tended to penalise minor syntactic or typographic errors that human instructors would overlook [116]. In rubric-based scoring of open-ended responses, AI sometimes failed to recognise algebraically equivalent or alternative correct solutions [43], and in handwritten answer grading, OCR errors and transcription failures led to missed or incorrectly evaluated responses [170]. In clinical assessment, AI graders using text-only inputs missed nuanced contextually appropriate communication [121], and in radiology report evaluation, sensitivity for minor discrepancies was particularly low [9]. False negatives were also documented in automated scoring of creativity [10], multilingual short-answer scoring [21], and reflective writing assessment [146].
Bias or fairness across student subgroups was evaluated in a smaller but growing number of studies, with findings ranging from confirmed bias to partial mitigation. The most thorough empirical subgroup analysis in the literature found that a DeBERTa-based scoring system showed only a negligible statistically significant effect for one demographic group (American Indian/Alaska Native, coefficient −0.02, marginal R² ≈ 0.0003), leading authors to conclude no notable differential performance [14]. In contrast, a study of LLM-based oral English assessment documented false positive rates varying by L1 background, with Arabic-speaking students showing 16% false positive rates before targeted mitigations [143]. Cultural and linguistic bias was identified as a concern in Japanese-language clinical communication assessment, where cultural response styles and tokenisation issues were acknowledged as potential sources of differential scoring [121]. A benchmark for Chinese educational values found that Chinese LLMs consistently outperformed English LLMs, confirming cultural bias as a measurable risk [73]. Studies of reflective writing assessment found that AI privileged polished English and formal style, raising concerns about disadvantaging non-native or non-standard writers [146], [209]. A study of LLM grading of nursing pitch scripts explicitly identified Western-centric pretraining as a potential source of bias against Hong Kong and Chinese students [161]. Country- and language-specific patterns in multilingual short-answer scoring suggested potential fairness issues across national groups [21]. A study of proficiency-stratified fairness in reflective writing found larger grading errors for low-ability learners [211]. Several studies acknowledged bias as a concern without conducting empirical subgroup analyses, calling for future fairness audits [123], [144], [169].
Gaming or adversarial manipulation of the AI grader was evaluated in fewer studies but confirmed as a real risk in several. A study of automated essay scoring found that AI graders failed to penalise incorrect or irrelevant content, and authors explicitly noted that students could exploit this tendency to receive undeserved credit by padding responses with off-topic material [54]. A study of open-ended mathematics scoring similarly observed that AI sometimes awarded full credit to inflated or off-topic responses and warned that students could exploit this [34]. A study of ChatGPT grading programming assignments noted that students could use ChatGPT to generate solutions, undermining the ability to assess genuine competence [46]. A bioinformatics course deployment implemented anti-cheating measures in system prompts and observed no prompt-hacking attempts during the semester, though authors cautioned that the risk remained [136]. A study of AI text detection tools documented students becoming familiar with circumvention strategies [193]. An educational values benchmark included adversarial prompts in its test set and found that retrieval-augmented generation substantially improved alignment on these challenging items [73]. A framework for AI-assisted assessment noted the possibility of prompt injection as an attack vector and suggested detection tools [49]. A study of an AI-based human-in-the-loop framework identified gaming and cheating detection challenges as a concrete concern, with 81% of surveyed respondents reporting noticing AI-generated content in student work [134]. A programming analytics framework detected anomalous submission behaviours consistent with copy-paste and possible external AI assistance [229].
Over-reliance or automation bias in human evaluators was evaluated in numerous studies, though typically as a concern raised and addressed through design recommendations rather than through direct empirical measurement of evaluator behaviour. Several studies documented the risk and recommended human-in-the-loop workflows as mitigations. A study of AI-assisted code review found that students tended to trust ChatGPT’s interpretations and relied on it for explanations, with authors recommending that students validate AI feedback [5]. A study of LLM grading of SOAP notes found that student trust in AI increased after disclosure, raising concerns about over-reliance [169]. A study of AI-assisted oral assessment found that learners could miss deviations in AI-generated patient responses and recommended supervised use and faculty review [184]. A study of AI-supported formative assessment observed student over-reliance on AI hints and noted the need for pedagogical safeguards [223]. A study of AI-assisted grading of lesson plans cautioned against replacing expert judgment and recommended human oversight [152]. A study of LLM grading in an astronomy course found that student over-reliance decreased over the semester as documentation and reflective requirements fostered verification workflows [215]. A study of AI-assisted oral performance assessment explicitly warned about the risk of over-reliance reducing teacher autonomy [224]. A study of GenAI-based grading of writing found that efficiency incentives could lead to over-reliance and that accountability was diffuse [209]. A study of AI grading of student reflections warned about “metacognitive laziness” and advised educators not to overly rely on GenAI insights [228].
Construct-irrelevant scoring — where AI graders rely on surface features such as length, keywords, or formatting rather than the intended construct — was confirmed as present in a large number of studies. AI graders were repeatedly found to emphasise surface-level linguistic features over deeper content quality. In essay scoring, ChatGPT was found to focus primarily on grammatical and spelling issues and to give generic, repetitive comments rather than substantive feedback [36], and to award higher scores to long and technical texts without adequately considering argument strength or logical coherence [155]. In reflective writing assessment, AI tended to emphasise surface textual features such as grammar and style, introducing construct-irrelevant variance tied to language proficiency rather than reflective ability [146]. In programming assessment, AI graders sometimes penalised superficial issues such as syntax errors and typos over student intent [116], and in code review, AI tended to identify more documentation and trivial issues while missing algorithmic efficiency and broader project trade-offs [5]. In short-answer scoring, bag-of-words approaches led to construct-irrelevant errors by missing synonyms and scoring based on numbers only [21]. In clinical assessment, AI systems emphasised formal linguistic features such as politeness and syntactic patterns over grounded clinical appropriateness [121], and in dental radiographic assessment, text-only LLMs could not process radiographs and may have missed image-dependent features [157]. In summary assessment, AI had difficulty distinguishing lexical copying from structural paraphrasing [147]. In user story evaluation, AI showed strong alignment on syntactic criteria but diverged on semantic and context-dependent criteria [206]. Several studies found that construct-irrelevant scoring could be partially mitigated through rubric alignment, coherence features, or knowledge-graph grounding [179], [166], [140].
Lack of transparency or explainability was formally evaluated in many studies and consistently confirmed as a concern. Authors across multiple domains noted that AI grading systems operate as black boxes, with scoring decisions that are difficult to interpret or audit. A study of GPT-4 coherence scoring noted that generated rationales may not reflect the actual rating process, so that explainability remains limited even when textual justifications are provided [13]. A study of LLM-based creativity scoring explicitly noted that LLMs can operate as black boxes and that explanation techniques may not scale [10]. A study of ChatGPT grading of dental radiographs found that students questioned the depth and interpretability of AI feedback and recommended integrating SHAP or LIME [157]. A study of GPT-4 grading of medical students’ OSCE performance noted the black-box problem and recommended explainable AI and uncertainty quantification [44]. A study of AI grading of SOAP notes found that the black-box nature of LLMs limited auditability [169]. In contrast, several studies found that transparency could be partially mitigated through rubric-in-prompt approaches [129], structured evidence-cited feedback [162], SHAP-based feature importance [12], or natural-language explanation generation [27], [140]. A study of AI-assisted grading found that natural-language insights were rated as more comprehensible and useful than word-level attribution highlights [228].
Inconsistency or unreliability across repeated runs was one of the most widely evaluated risks in the literature and was confirmed as present in the majority of studies that tested it. Multiple studies documented that identical inputs submitted to the same model on different occasions or in different sessions produced different scores. A study of ChatGPT grading of programming assignments found that multiple grading attempts on identical inputs produced markedly different results, concluding that this inconsistency disqualified ChatGPT as a standalone assessor [46]. A study of GPT-4 grading of calculus exercises documented randomness and temporal fluctuations, low correlations across runs, and degradation of coherence over long sequences [43]. A study of ChatGPT scoring of reflective essays observed inconsistent scores across three repeated ratings of the same essay (91, 93, 88) [20]. A study of GPT-4 grading of short-answer questions found that Gemini showed higher variability than GPT-4 across ten repeated runs [49]. A study of GPT-4 grading of physics solutions found that some models showed large variation across repeated trials linked to chatbot randomness and temperature settings [61]. A study of LLM grading of programming assignments found substantial differences between models and variants, with some models dominated by failing grades and others by maximum scores [117]. A study of Chinese high school essay grading found that DeepSeek showed score instability with a maximum difference of five points between two runs [165]. Several studies found that inconsistency could be mitigated through temperature control [26], averaging across multiple runs [129], ensemble approaches [14], or customised GPT configurations [181]. A study of GPT-4 grading of student responses found very high short-term consistency (ICC = 0.999) but lower consistency after several weeks (ICC = 0.944), indicating temporal drift as a distinct reliability concern [48].
Privacy or data security was evaluated in a smaller but meaningful subset of studies. Several studies identified the use of commercial cloud-based LLM APIs as a privacy risk because student data are transmitted to external servers. A study of radiology report evaluation recommended open-source on-premise models for institutional control due to data ownership and privacy concerns [9]. A study of automated essay scoring identified privacy concerns with remote commercial generative models and mitigated them by running the system locally [114]. A study of LLM grading of programming assignments implemented pseudonymisation, IRB protocols, and OpenAI enterprise API assurances and cited GDPR and FERPA compliance as design drivers [169]. A study of automated formative assessment ran the system locally using Ollama, anonymised student identities, and stored data on encrypted university systems [138]. A study of medical learning log extraction explicitly noted that many LLMs are cloud-based and called for new approaches to address privacy concerns [124]. A study of AI-assisted grading of student writing identified ethical concerns about uploading student work to AI systems and called for explicit data privacy and consent protocols [170]. A study of LLM grading in a university context noted that reliance on ChatGPT, a third-party commercial product, raised concerns about governance and data control [10]. A study of OSCE video segmentation used HIPAA- and FERPA-compliant APIs for proprietary models and hosted open-source alternatives on secure institutional infrastructure [204]. A study of AI-based assessment in Norwegian medical education recommended compliance with GDPR, the AI Act, and national regulations [44].
Beyond the ten canonical risk categories, several studies formally evaluated additional risks not captured by the standard taxonomy. Hallucination — the generation of plausible but factually incorrect content — was evaluated and confirmed in multiple studies. A study of ChatGPT grading of calculus exercises documented hallucinations including fabricated rows and incorrect algebraic steps as principal causes of grading errors [43]. A study of physics solution grading identified mathematical hallucinations as a repeated source of over-leniency [61]. A study of an LLM-based clinical oral examination simulator measured a 4.7% hallucination rate and implemented detection and routing to human reviewers [166]. A study of an anesthesiology SAQ grader measured 67 hallucinations across 316 AI-generated patient responses [184]. A study of ChatGPT grading of academic essays found that the model fabricated non-existent or incorrect references [20]. A study of an LLM-based CPX chatbot acknowledged that the inherent possibility of hallucination could not be completely eliminated [26]. Anchoring bias — where AI grading is influenced by prior knowledge of human grades — was formally evaluated in one study, which found that providing the human grade to the LLM before scoring introduced a systematic bias [49]. Systematic scoring biases including central tendency and stringency effects were documented in a study of GenAI assessment of MCQ cognitive levels, where AI showed systematic stringency in Miller classifications and central tendency bias in Bloom classifications [205]. The risk of AI grading producing consequential shifts in grade distributions was evaluated in a study of SOAP note grading, which found that ChatGPT produced a statistically significant increase in the honours rate [178]. Modality limitations — the inability of text-only AI systems to evaluate visual, auditory, or multimodal content — were confirmed as a risk in studies of OSCE video assessment [47], radiographic evaluation [157], and surgical video quality assessment [177]. The risk of AI grading systems replicating shared biases through high inter-model consensus was identified in a study comparing 18 LLMs as automated graders, which warned that high agreement across models could mask systematic errors [117]. The risk of AI text detection tools generating false accusations against students — a form of false positive with direct disciplinary consequences — was formally evaluated and confirmed in a study of AI detection in higher education [193].
Remediation Strategies
Where studies moved beyond documenting risks to actively testing countermeasures, several coherent families of remediation strategy emerge from the literature. These are best understood by the risk they target, the mechanism of intervention, and the degree to which they resolved the problem.
Prompt engineering and rubric-aligned prompting is the most widely tested single class of mitigation, applied across a large number of studies targeting low agreement, construct-irrelevant scoring, and inconsistency. The core logic is that providing the model with explicit scoring criteria, structured output formats, and worked examples constrains its generative freedom and anchors judgments to pedagogically meaningful constructs. The effect is consistently positive but conditional. Studies of short-answer and essay scoring found that rubric-guided prompting substantially improved agreement with human raters compared to unguided or zero-shot conditions: one study reported that unguided AI scoring produced very low agreement (ICC = 0.076) while structured prompts raised this to ICC = 0.708 [152], and another showed that rubric provision in the prompt transformed near-chance agreement into high alignment with expert scores [34]. Similar gains were reported in physics grading, where supplying a mark scheme substantially reduced leniency and improved score distributions [61], and in dental education, where including the rubric in the prompt improved AI-student alignment in more than half of cases [25]. In programming assessment, rubric-dependent experiments showed that scoring varied substantially with rubric emphasis, confirming that rubric clarity is a prerequisite for construct-valid scoring [169]. Prompt engineering also proved effective for reducing construct-irrelevant scoring in a study of student-story evaluation, where rubric-guided prompts produced high agreement on structural criteria while revealing persistent gaps on semantic ones [206]. However, rubric provision did not universally improve performance: one medical education study found that supplying expert rubrics improved some model-question combinations while worsening others, and that LLMs sometimes misapplied rubric criteria in ways that introduced new errors [175]. A study of music education similarly found that customized GPTs with embedded rubrics produced more stable scoring than repeated standard prompts, but that domain-specific knowledge gaps remained [181].
Few-shot prompting and example inclusion is closely related and frequently tested alongside rubric provision. Providing graded exemplars in the prompt improved classification accuracy for short-answer scoring [136], reduced variability in text-complexity classification [183], and improved match rates in reflection scoring [213]. A study of programming feedback found that few-shot prompting reduced variance relative to zero-shot conditions, though it did not eliminate inconsistency [27]. In one study, even a single example substantially improved performance over zero-shot baselines [202]. The choice of examples matters: semantic-search-based example selection reduced variability compared to random sampling [183], and using examples aligned with the target evaluator’s perspective improved generalization [183]. Conversely, few-shot prompting for physics concept classification did not improve ChatGPT performance and in one case reduced F1, illustrating that the strategy is not universally effective [24].
Repeated runs and score aggregation address the well-documented nondeterminism of LLM outputs. The most common implementation involves running the same prompt multiple times and averaging or taking the median of the resulting scores. This approach was tested in studies of essay grading [2], programming assessment [169], short-answer scoring [129], and knowledge-game scoring [26], with consistent findings that aggregation reduces variance and improves reliability relative to single-run outputs. One study reported that averaging five runs produced high test-retest reliability (Cronbach’s α = 0.98) [182], and another found that five independent runs of an AI behaviour assessment tool yielded ICC values of 0.94–0.96 when averaged, substantially exceeding single-run human interrater reliability [123]. Setting temperature to zero was tested as a complementary approach to reduce stochasticity, with one study reporting that this produced consistent outputs for a knowledge-game scoring system [26] and another using it to enable reproducibility in a clinical OSCE study [121]. However, even temperature-zero settings do not guarantee full determinism, as one study noted that GPT-4 behaviour can shift over time regardless of temperature [11].
Fine-tuning and task-specific model training is the most resource-intensive mitigation but consistently produces the largest accuracy gains when training data are available. A radiology education study found that a task-specific fine-tuned model (RadBERT) substantially outperformed zero-shot generative LLMs on discrepancy detection [9]. Fine-tuning GPT-3.5 on annotated feedback quality data improved F1 from around 0.49 to 0.83 on the most difficult dimension [196]. A self-augmentation procedure for BERT-based models improved performance on small educational datasets across two languages and two tasks [45]. In a translation quality assessment context, fine-tuning LLMs on domain-specific data reduced hallucination and improved fluency and faithfulness [159]. A study of short-answer grading found that a fine-tuned smaller model (AERA, distilled from ChatGPT outputs) improved QWK by eleven percentage points over the base ChatGPT and produced more consistent rationales [27]. Domain-specific fine-tuning also improved pronunciation assessment, with targeted data augmentation for underrepresented language groups reducing false positive rates [143]. The limitation is that fine-tuning requires annotated data, computational resources, and periodic retraining as models and curricula evolve.
Ensemble methods and hybrid architectures combine multiple models or model types to achieve performance that no single component reaches alone. A study of classroom observation scoring found that a weighted ensemble combining a supervised MLP regressor with GPT-4 achieved human-level interrater reliability (r = 0.513), whereas each component alone fell short [12]. A multi-agent architecture (AutoSCORE) that decomposed scoring into a rubric-component extraction agent followed by a scoring agent improved agreement and interpretability compared to single-agent baselines [140]. A hybrid rule-based and LLM architecture (HARRS) for report assessment achieved MAE of 0.73 and Pearson r of 0.98 against human scores, compared to MAE of 5.69 and r of 0.46 for an unguided LLM baseline [139]. A collaborative multi-agent system (CAMS) using few-shot chain-of-thought prompting and a collaborative module improved accuracy and reduced stochastic instability for small locally deployed models [145]. A RAG-enhanced evaluation system for teacher professional competency improved performance on ethics and professional philosophy dimensions and showed robustness to adversarial prompts [73].
Score calibration and post-processing targets systematic bias in AI scores, particularly leniency or stringency relative to human standards. One study tested regression-based mapping of LLM scores to human-equivalent letter grades to correct for systematic stringency, finding that while rank ordering was preserved, absolute score differences required correction [215]. Another study applied iterative hyperparameter calibration of scoring weights and tolerance thresholds, correcting an initial five-point bias and achieving ICC of approximately 0.95 [138]. A study of NAEP math scoring used data augmentation with a generative LLM to balance underrepresented score classes and applied ensemble cross-validation to improve stability [14]. Statistical rescaling of AI grades was tested in a physics grading context but found insufficient to correct dispersion or feedback mismatches unless the base correlation was already high [61]. A study of creativity scoring used multiple aggregation methods for originality scores and per-session normalization to address score inflation and distributional biases [10]. In essay grading, applying SMOTE to balance imbalanced score classes and adding coherence features raised QWK from 0.69 to 0.89, approaching human interrater consistency [179].
Human-in-the-loop review thresholds and hybrid workflows represent the most broadly recommended structural mitigation across the literature, though the specific implementation varies. The most common design positions AI as a first-pass scorer or triage tool, with human review triggered by disagreement, low confidence, or flagged cases. A study of NAEP scoring recommended that the model replace one human rater, with a second human invoked only when the model and remaining human disagree [14]. A regrade request policy allowing students to request human review was implemented in a large-scale university deployment and resulted in low regrade rates, suggesting the AI-human hybrid was acceptable to students [120]. A sequential tolerance-based triage system automatically accepted high-confidence scores, re-evaluated borderline cases, and routed persistent disagreements to mandatory human review [138]. A study of clinical history-taking assessment used a multi-stage optimization pipeline with original-text matching and keyword checks to prevent unwarranted scoring, achieving ICC above 0.923 [162]. A study of OSCE video analysis implemented human review of flagged hallucinations, adding approximately two hours per week of oversight at experimental scale [166]. In a short-answer grading system, mandatory faculty review was built into the design, with the system functioning as a recommender rather than an autonomous grader, and this design choice was credited with reducing faculty concerns about over-reliance [114]. A study of programming assignment grading embedded anti-cheating measures in system prompts and provided a student appeal option for manual review, reporting no prompt-hacking attempts during the semester [136].
Bias detection and fairness-aware evaluation was formally tested in a smaller number of studies. A study of NAEP math scoring conducted pairwise standardized mean difference analyses and linear mixed-effects regression to detect demographic prediction error, finding no meaningful bias overall though one small statistically significant effect was observed for one subgroup [14]. A pronunciation assessment study conducted cross-cultural bias analysis and implemented targeted mitigations including multi-accent training data augmentation, adaptive assessment thresholds by first language, and teacher bias-awareness training, reporting residual false positive rates of 16% for Arabic-speaking learners before adjustments [143]. A multi-agent feedback system implemented an Equity Monitor agent to check for biased or exclusionary language in AI-generated comments, and used stratified MAE by proficiency band to make error disparities explicit [211]. A study of OSCE scoring conducted subgroup and sensitivity analyses across station type, year, prior clerkship, and institution, finding no significant differences in AI performance across these groups [163]. A study of radiology discrepancy detection evaluated pathology-type distribution in the AI-curated set and found no pathology-type bias, though demographic learner-subgroup fairness was not assessed [9].
Transparency and explainability interventions were tested primarily through the provision of rationales alongside scores. A study of short-answer grading found that requiring the model to generate rationales improved interpretability, and that human evaluation showed AERA rationales were comparable or preferred to ChatGPT’s, though the authors cautioned that rationales may not reflect the actual scoring process [27]. A study of classroom observation scoring applied SHAP to a supervised model and found that feature importance aligned with rubric-relevant constructs, providing evidence against construct-irrelevant scoring [12]. A study comparing important-word highlights to natural-language rationales found that natural-language explanations were rated as more comprehensible and useful by educators, and that word-level highlights sometimes encouraged surface strategies [228]. A multi-stage clinical assessment pipeline generated verbatim dialog citations for every scored item, achieving 100% compliance with predefined transparency criteria [162]. A hybrid architecture explicitly separated auditable deterministic rules from probabilistic LLM outputs to avoid black-box assessment [139]. Transparency was also pursued through open-source code release and detailed system-message documentation [55].
Privacy-preserving deployment strategies were tested in a smaller set of studies. Running models locally rather than via commercial APIs was implemented in a short-answer grading system to prevent student data from reaching third parties [114], and recommended in a radiology education study as a mitigation for data ownership concerns [9]. De-identification of student data prior to model input was implemented in OSCE grading [115] and radiology report evaluation [31]. A model-size reduction strategy identified a much smaller model (BERT-mini) that matched baseline performance, enabling potential on-device deployment to enhance privacy [41]. A programming assessment study implemented pseudonymization, used enterprise API agreements, and documented GDPR compliance as design drivers [169]. A clinical simulation study used the publicly available anonymized MIMIC-IV dataset and noted IRB exemption, addressing privacy through data source choice [102].
Across all these strategies, a consistent pattern emerges: no single mitigation is sufficient in isolation, and the most effective implementations combine multiple approaches — typically rubric-aligned prompting, repeated-run aggregation, and human-in-the-loop review — while the most ambitious studies add fine-tuning, ensemble methods, or formal bias audits. Strategies that were effective in one context frequently showed conditional or null effects in others, underscoring the importance of task-specific validation before deployment.
Limitations
The limitations acknowledged across this body of empirical literature cluster into five recurring themes, each reflecting a distinct vulnerability in the current evidence base for AI-based educational evaluation.
Sample and Scope
The most pervasive limitation is the narrowness of the samples and contexts from which findings are drawn. Single-institution designs dominate the literature, with studies conducted at one university, in one course, or with one cohort, a pattern acknowledged explicitly by authors across medical, engineering, language, and science education contexts [54], [129], [163], [62], [138]. Participant counts are frequently small — sometimes fewer than fifteen — and authors routinely note that this restricts statistical power and external validity [36], [102], [123], [56], [208]. Task coverage is similarly constrained: many studies examine only one or two assignment types, such as a single reflective essay and a coding task [2], two programming assignments [3], or a single clinical scenario [26], [102]. Linguistic and demographic coverage is equally narrow: studies conducted in Japanese, Indonesian, Hebrew, or Chinese explicitly acknowledge that findings may not transfer across languages or cultural contexts [121], [206], [10], [73], while studies relying on non-native speaker populations note that model performance degrades for learners whose language differs from dominant training corpora [143], [161].
Evaluation Methodology
Methodological limitations are widespread and often compound one another. A recurring problem is the use of a single human rater as the reference standard, which precludes assessment of inter-human variability and conflates human idiosyncrasy with model error [49], [178], [157], [184]. Several studies compare AI output against a consensus of only two or three experts, which, while more robust than a single rater, still represents a narrow gold standard [175]. Evaluation windows are typically short and cross-sectional, capturing model behaviour at a single point in time without assessing stability across sessions or model updates [156], [187], [doi:10.30774/wjaets.2026.18.3.0182]. Prompting strategies are rarely varied systematically: many studies use a single prompt formulation or a zero-shot approach and acknowledge that alternative prompt designs might yield substantially different results [144], [152], [200], [214]. Reliance on a single metric type — most commonly correlation or agreement coefficients — without complementary fairness or validity analyses is also noted as a methodological gap [179], [229]. In multimodal assessment contexts, the absence of formal annotation reliability studies further weakens the evidentiary base [204].
Generalisability
Even where studies are internally valid, authors frequently caution that findings were obtained under controlled or low-stakes conditions that do not reflect authentic high-stakes deployment. Laboratory-derived datasets, pilot designs, and proof-of-concept implementations are common, and authors acknowledge that performance under real classroom conditions — with diverse student populations, genuine assessment stakes, and institutional infrastructure constraints — remains untested [11], [134], [163]. Longitudinal follow-up is almost entirely absent: the vast majority of studies capture a snapshot of AI grading performance without examining whether agreement, reliability, or student outcomes persist over time [128], [203], [211]. Authors studying AI feedback systems similarly note that downstream learning effects were not measured, so efficiency or agreement gains cannot be interpreted as evidence of educational benefit [120], [138], [168]. The controlled conditions of several studies — including scripted clinical encounters, pre-selected essay samples, and constrained assignment formats — further limit the ecological validity of reported findings [163], [102], [218].
AI-Specific Limitations
A cluster of limitations is specific to the properties of large language models as evaluation instruments. Prompt sensitivity — the finding that small changes in prompt wording materially alter scoring behaviour — is acknowledged across diverse task types and models [2], [3], [136], [169], [186]. Model versioning instability compounds this: because LLMs are continuously updated, results obtained with one version may not replicate with a later release, and several authors note that the specific model version used in their study is no longer available [222], [11], [187], [43]. The black-box opacity of model scoring processes is a widely acknowledged concern: authors note that generated rationales may not reflect the true decision process, that explanations are difficult to audit, and that this opacity undermines trust and accountability in assessment contexts [10], [13], [20], [44], [48]. Dependency on training data characteristics manifests in multiple ways: models trained predominantly on English-language or Western academic text show degraded performance on non-dominant languages, non-standard student phrasing, and domain-specific content that is underrepresented in pretraining corpora [24], [40], [129]. Nondeterminism — the production of different outputs for identical inputs across runs — is acknowledged as a reproducibility threat even when temperature is set to zero [11], [121], [136]. Multimodal limitations are also prominent: text-only models cannot evaluate handwritten work, circuit diagrams, radiographic images, nonverbal communication, or other non-textual assessment artefacts, a constraint that substantially narrows the range of educational tasks to which current systems can be applied [142], [222], [157], [141].
Missing Evidence
A final category of limitation concerns the systematic absence of evidence that authors themselves flag as gaps in their own studies. Accuracy is frequently assessed without any accompanying fairness analysis: studies report agreement coefficients or error rates without examining whether performance is equitable across student subgroups defined by language background, prior achievement, or demographic characteristics [146], [123], [179]. Efficiency gains — reduced grading time, increased throughput — are reported without corresponding data on whether the quality of student learning is maintained or improved [138], [168], [120]. Risks are frequently acknowledged in limitations sections without formal evaluation: authors note the possibility of hallucination, scoring bias, or adversarial manipulation but do not empirically test the frequency or severity of these phenomena [54], [136], [37]. The quality of AI-generated explanations and feedback — as distinct from the accuracy of scores — is rarely assessed, leaving open the question of whether feedback that accompanies correct scores is itself pedagogically sound [156], [197], [27]. Taken together, these gaps mean that the literature provides a partial picture: it documents what AI systems can score, but offers limited evidence about whether doing so is fair, durable, or educationally beneficial.
Recommendations
The recommendations literature converges on four interconnected themes that together constitute a framework for responsible deployment of AI in assessment evaluation: maintaining meaningful human oversight, addressing bias and fairness proactively, ensuring transparency and explainability, and managing institutional integration carefully.
Human-in-the-Loop Design
The most consistent recommendation across the literature is that AI should function as an assistant or first-pass tool rather than an autonomous decision-maker. Virtually every empirical study that tested AI grading systems concludes with some variant of this principle. Authors recommend using AI to reduce the volume of work reaching human reviewers — flagging borderline cases, triaging large submission sets, or generating draft feedback — while reserving final evaluative authority for instructors [2], [3], [11], [46]. This position is not merely precautionary: several authors articulate specific triage architectures in which AI identifies the cases most warranting human attention, such as using an LLM to flag score discrepancies for educator review [9] or routing low-confidence and borderline responses to instructors while accepting AI scores on clear-cut cases [138], [144].
The question of where to set the automation threshold is treated with considerable nuance. Several authors recommend restricting fully automated scoring to formative, low-stakes contexts and requiring human review for summative or high-stakes decisions [55], [49], [175], [178]. A phased deployment model is widely endorsed: begin with formative triage, move to hybrid summative grading with human audit of flagged responses, and only consider selective automation of clearly objective items once reliability has been established [144], [187]. The recommendation to provide students with an explicit right to request human review appears in multiple studies as a structural safeguard against automation bias [136], [182], [146].
To preserve efficiency gains while preventing over-reliance, authors recommend several practical mechanisms. Repeating AI scoring runs and averaging results reduces stochastic variability without requiring additional human effort [2], [61], [116]. Ensemble approaches — combining multiple models or aggregating LLM scores with human ratings — are recommended to reduce individual model idiosyncrasies [16], [186]. Authors also caution against presenting AI scores to human reviewers before they have formed their own judgments, since doing so risks anchoring bias [49]. The broader principle is that human oversight should be structurally embedded in the workflow rather than treated as an optional fallback.
Bias and Fairness
Recommendations for bias detection and mitigation cluster around three stages: pre-deployment auditing, fairness-aware design, and ongoing monitoring. Before deployment, authors recommend conducting formal subgroup analyses to detect disparate impact across demographic groups, language backgrounds, and cultural contexts [123], [143], [144]. Pairwise standardized mean difference analyses and mixed-effects models are specifically recommended for detecting algorithmic bias by demographic group [14]. Several authors note that bias audits should be treated as mandatory rather than aspirational, with predefined procedures for pausing or adjusting AI scoring if disparities are detected [123].
At the design stage, recommendations include diversifying training data to represent varied expression styles, linguistic backgrounds, and cultural communication norms [138], [161], [143]. Multi-accent training data and culturally responsive feedback calibration are specifically recommended for spoken language assessment contexts [143]. Data augmentation techniques — including paraphrase generation and oversampling of underrepresented score categories — are recommended to address class imbalance that can systematically disadvantage certain student populations [14], [179]. Authors working in multilingual contexts recommend evaluating automated scoring not solely against human benchmark scores but also against explicit fairness criteria, since human raters themselves may carry cultural biases [21].
For ongoing monitoring, authors recommend periodic revalidation of AI scoring systems against human ratings, with particular attention to model drift over time [11], [49], [62]. Equity-focused monitoring agents embedded in the scoring pipeline — capable of comparing error rates across learner bands and triggering targeted calibration — are proposed as a structural mechanism for continuous fairness assurance [211]. The recommendation to instruct prompts to focus on content rather than length or stylistic features, and to diversify training data to accommodate varied writing styles, addresses a specific and frequently documented source of systematic unfairness [138].
Transparency and Explainability
A recurring recommendation is that AI scoring systems should produce not only scores but also rubric-aligned rationales that make the basis for each judgment visible to instructors and students. Requiring the AI to generate per-dimension explanations alongside scores — rather than a single holistic grade — is recommended both for quality control and for pedagogical value [194], [140], [13]. Several authors recommend structuring the scoring pipeline so that evidence extraction and scoring are performed as separate steps, creating an auditable chain from rubric criterion to evidence to judgment [140]. Rationale-first or rating-first prompt configurations that elicit explicit reasoning are recommended over opaque single-token outputs [13].
Communicating confidence levels is identified as a distinct transparency requirement. Authors recommend incorporating explicit confidence indicators — whether through probabilistic outputs, flagging of low-confidence cases, or uncertainty quantification — so that instructors can prioritize their review effort and students understand when a score is provisional [138], [44], [176]. Explainability methods such as SHAP and LIME are recommended for post-hoc inspection of model decisions, particularly in high-stakes contexts where instructors need to understand why a particular score was assigned [12], [157], [229].
Transparency obligations extend to students and institutional stakeholders. Authors recommend that institutions disclose to students when and how AI is involved in their assessment, what data are used, and what recourse is available [138], [209], [193]. The right to contest an AI-generated score through a human review process is framed not merely as a safeguard against errors but as a condition for maintaining student trust and institutional legitimacy [146], [182]. Several authors recommend that AI-generated feedback be explicitly labelled as such and that students be trained to evaluate it critically rather than accept it uncritically [173], [171].
Deployment and Institutional Integration
Authors consistently recommend piloting AI evaluation tools incrementally before broad deployment. A staged approach — beginning with a single course or assessment type, evaluating task-technology fit, and expanding only where fit is strong — is endorsed across multiple studies [187], [134], [60]. Pilot phases should include comparative evaluation against human grading, usability testing with instructors, and monitoring of student outcomes, with explicit go/no-go criteria before scaling [60], [218]. Several authors recommend that rubric-aligned prompts and standardized prompt templates be developed and validated during the pilot phase, since prompt design has been shown to have a large effect on scoring quality [129], [169], [203].
Faculty preparation is treated as a prerequisite rather than an afterthought. Recommendations include structured professional development covering prompt engineering, critical interpretation of AI outputs, and recognition of AI limitations [143], [186], [181]. Authors emphasize that faculty should retain ultimate responsibility for grades and that institutional policy should make this explicit rather than allowing ambiguity about the locus of evaluative authority [209], [138]. Student preparation — including AI literacy training and guidance on how to interpret and critically engage with AI feedback — is recommended as a parallel investment [56], [193].
Privacy and data governance recommendations are prominent throughout the literature. Authors recommend deploying models on-premises or using institutional enterprise agreements rather than consumer interfaces, anonymizing student data before submission to external APIs, obtaining informed consent, and ensuring compliance with applicable data protection regulations [9], [49], [138], [170]. The use of open-source models deployed locally is specifically recommended as a means of retaining institutional control over student data while avoiding dependence on commercial providers whose terms of service may be incompatible with educational privacy requirements [145], [202].
On the question of appropriate scope, the literature is notably cautious about extending AI evaluation authority to high-stakes summative decisions, assessments of complex interpersonal or clinical competencies, or contexts where cultural and linguistic diversity has not been adequately represented in model training [121], [163], [222]. The recommendation that AI decision-making authority be bounded by the demonstrable reliability of the system in the specific deployment context — rather than by general claims about model capability — reflects a broader principle of evidence-based, context-sensitive adoption that runs through the recommendations literature as a whole [117], [44].
Contributions from Non-Empirical Research
The non-empirical literature surveyed here spans systematic and scoping reviews, conceptual frameworks, design proposals, taxonomies, and policy-oriented guidance documents. Together these contributions provide the theoretical scaffolding, classificatory vocabulary, and normative principles that empirical studies often presuppose but rarely articulate — and they identify a substantial number of open research directions not yet addressed by the empirical record.
Taxonomies and Classification Frameworks
Several papers make their primary contribution by imposing classificatory order on a fragmented empirical landscape. The most architecturally ambitious of these proposes an Input-Process-Output (IPO) theoretical framework and a five-type taxonomy of text-based automated assessment systems — Automatic Grading, Automatic Classifier, Automated Feedback, Automated Writing Evaluation, and Multimodal Evaluation — synthesised from post-secondary literature [15]. This taxonomy provides a shared vocabulary that cuts across the domain-specific reviews discussed below. A complementary bibliometric mapping and systematic review constructs a macro-level publication landscape alongside a micro-level qualitative synthesis, classifying AI-assisted assessment tools by effectiveness, technical design, and ethical challenge, and identifying longitudinal validation and ethics as the most urgent research needs [113]. A PRISMA-guided systematic review of AI in physics education offers a thematic taxonomy of AI uses — instruction, tutoring, and automated assessment — and documents the recent dominance of LLMs in assessment research while aggregating study-level recommendations for human–AI collaborative grading [130]. For language education specifically, a comprehensive systematic review combining bibliometrics, inductive content analysis, and Structural Topic Modelling maps AI technologies including automated writing evaluation and automatic speech recognition across a decade of research, producing a topical taxonomy and identifying diversity, equity, and inclusion as a conspicuous gap [118]. A parallel review focused on foreign-language higher education provides a thematic classification of AI applications including automated assessment, synthesising advantages and risks and identifying gaps in standards and teacher preparation [131].
Domain-specific scoping and systematic reviews extend this classificatory work into professional and clinical education. A scoping review of AI in emergency medicine education proposes a taxonomy of educational applications — Evaluation and Feedback Systems, Skills Assessment, and Simulation — and synthesises evidence on LLM performance in exams and narrative comment scoring while flagging the absence of external validation and interpretability work [50]. A parallel scoping review of AI in orthopaedic surgical training identifies two core application themes — refinement of surgical competencies and enhancement of knowledge acquisition — and catalogues AI models used, including LLMs, while foregrounding personalised automated feedback as a promising but undervalidated direction [35]. A synthesis focused on graduate medical education maps NLP and LLM applications to automated performance evaluation, sentiment analysis of narrative feedback, personalised learning recommendations, and competency-assessment algorithms, while emphasising the absence of multi-institutional benchmarking [59]. Two systematic reviews of ChatGPT in dental education provide overlapping thematic frameworks — one organising findings across perceptions, educational uses, benefits, risks, curricular integration, and model comparison [58], the other synthesising research trends and study quality while classifying a subset of studies on feedback and grading [32]. A systematic review of AI in periodontal education synthesises psychometric analyses of AI-generated exams and comparisons of LLM grading against human grading, identifying automation bias and variable LLM reliability as recurring risks [127]. A survey of LLM applications in Traditional Chinese Medicine education identifies data scarcity, model adaptability, and the absence of standardised evaluation frameworks as field-specific challenges and proposes a TCM-specific evaluation framework as a design priority [83]. A PRISMA-guided systematic review of ChatGPT across forty studies in general education synthesises opportunities and drawbacks including automated grading, maps governance and policy challenges, and highlights the limited real-world validation of automated grading systems [70]. A systematic review of GPT-3 in education synthesises thirty-four studies covering question generation, evaluation of student-generated items, and automated diagnostics, documenting performance comparisons and limitations including bias and context constraints [92]. A systematic review of generative AI in K-12 education applies PRISMA and the Mentefacto Map to identify a notable absence of concrete empirical studies and longitudinal research, and proposes frameworks, teacher training, and ethical guidelines as priorities [80]. A systematic review of AI integration in computer-based testing consolidates multidisciplinary evidence, maps the rise of generative AI and adaptive testing, quantifies frequencies of applications and risks, and identifies the absence of explainable AI, inclusivity provisions, and blockchain-based security as research gaps [189].
Conceptual Frameworks for AI-Integrated Assessment Design
Beyond classification, a second cluster of non-empirical contributions proposes design frameworks that specify how AI should be integrated into assessment workflows. A conceptual roadmap for integrating AI with learning progressions and constructed-response assessments articulates staged uses of AI across LP development, validation, scoring, feedback, and instructional support, and maps specific AI approaches to LP stages while foregrounding measurement validity and equity as design constraints [30]. This framework directly extends the empirical literature on automated scoring by providing principled criteria for when and how AI scoring should be introduced. A conceptual framework for AI-based second-language speaking assessment articulates requirements for validity, reliability, fairness, practicality, and ethics, surveys existing tools against those requirements, and proposes a combined framework and workflow drawing on educational design research [160]. A conceptual synthesis for automated writing evaluation in EMI and ESL contexts differentiates AWE, AES, and GenAI, proposes a hybrid Human–AI feedback model with a sample rubric and task-distribution table, and identifies longitudinal and cross-cultural validation as open research directions [137]. A position paper on LLMs in language teaching and automated assessment surveys state-of-the-art findings and proposes practical principles including calibration, explainability, and human-in-the-loop deployment [51]. A conceptual synthesis mapping IGO guidelines to learning design and assessment consolidates competencies in higher-order thinking and AI literacy alongside ethical principles of privacy, transparency, and equity, and proposes AI-inclusive assessment design as a practical direction [81]. A methodological proposal for using generative AI dialogues to evaluate pre-service science teachers’ pedagogical content knowledge outlines a two-stage method — recording teacher–GenAI lesson-planning conversations and assigning trained GenAI to analyse them — and situates the approach within PCK and TPACK frameworks while emphasising training and validation needs [8]. A six-step framework for developing customised GPT models in medical education operationalises constructivist educational theory, catalogues fifteen prototype GPTs including assessment-related tools, and offers implementation guidance [94]. A consensus-driven set of twelve tips for designing take-home and research-based assessments in the generative AI era synthesises practitioner perspectives into four design principles — co-developing AI literacy, prioritising process over output, validation and verification strategies, and preparing students for AI-enhanced workplaces [87]. A conceptual analysis of LLMs for residency training assessment in China proposes an ethics- and governance-oriented framework covering data curation, privacy, bias mitigation, human oversight, and accountability [104]. A narrative synthesis for family medicine education proposes a tiered faculty competency model — Basic, Proficient, and Expert — to guide workforce development for using and governing AI-based assessment, alongside recommendations for human-in-the-loop deployment, algorithmic auditing, and governance structures [226].
Governance, Ethics, and Risk Frameworks
A third cluster addresses the normative and governance dimensions of AI-based assessment. The most structurally comprehensive of these contributions is an AI Risk Management Framework that proposes four governance functions — Govern, Map, Measure, and Manage — and a taxonomy of trustworthiness characteristics including validity and reliability, safety and security, accountability and transparency, explainability and interpretability, privacy enhancement, and fairness with managed bias [225]. This framework is notable for being explicitly cross-cutting: it applies to any AI system used in evaluative roles and provides a vocabulary that connects the empirical findings on bias and accuracy failures documented elsewhere in this catalog to a structured governance response. A conceptual review of generative AI in medical education applies an ecological framing across micro, meso, and macro levels of regulation to map pedagogical impacts, trustworthiness concerns, and academic integrity issues, and proposes assessment redesign and governance as priorities [93]. A conceptual synthesis of GenAI in education highlights ethical challenges and academic integrity risks and proposes design principles including open and ethical LLMs and interdisciplinary co-creation [42]. A conceptual mapping of LLM applications across the test development and assessment cycle — item generation, administration, scoring, analysis, and reporting — emphasises the collaborative role of teachers with AI and highlights bias, privacy, and oversight as limitations requiring practical strategies [39]. A qualitative literature-review-based synthesis for Islamic Religious Education proposes a normative framework grounded in Islamic principles of public benefit, justice, and trustworthiness, identifies domains where AI can assist across cognitive, affective, and psychomotor dimensions, and catalogues ethical and pedagogical challenges including the limits of AI in assessing spirituality [201]. An essay on AI in nursing education provides policy- and practice-focused recommendations covering privacy, bias, and academic integrity [64], while a parallel synthesis for nursing education summarises current uses of AI-generated NCLEX-style questions and automated feedback, calls for guidelines and faculty development, and identifies the absence of empirical evaluation of AI interventions as a gap [97]. A guidance piece on ChatGPT in vocational and professional education maps capabilities, limitations, and implementation recommendations including human oversight and multimodal outputs [98].
Design Tools and Platform Proposals
Two contributions occupy a distinct category by proposing concrete computational artefacts rather than frameworks alone. A systems paper introduces Open Brain AI, a platform integrating machine learning, NLP, LLMs, and automated speech-to-text to compute multilingual linguistic metrics at discourse and granular levels, offering a practical tool to automate language assessment workflows for clinicians, researchers, and educators [52]. A qualitative analysis of how ChatGPT-3.5 characterises digital learning assessment proposes an illustrative evaluative activity model that positions AI as an object and tool for learning rather than as an autonomous grader, contributing a pedagogical design principle that complements the empirical literature on student-facing AI feedback [75].
Open Research Directions
Across these contributions, several themes recur as open directions not yet adequately addressed by the empirical record. The absence of standardised evaluation criteria and methodological heterogeneity is flagged by multiple domain-specific reviews [32], [58], [59], making cross-study comparison difficult and limiting the generalisability of accuracy findings. Longitudinal research is consistently identified as absent [80], [118], [113]. Explainability and interpretability — foregrounded in the AI RMF [225] and the emergency medicine review [50] — remain largely unaddressed in empirical deployments. Fairness and equity, including diversity and inclusion considerations in automated scoring, are identified as gaps in language education research [118] and as design requirements in speaking assessment frameworks [160]. The governance of AI in assessment — including human oversight structures, accountability mechanisms, and faculty competency development — is proposed by multiple non-empirical contributions [226], [225], [104] but has not yet been studied empirically. These convergent gaps constitute a clear agenda for the next generation of empirical work.
Conclusion and Future Directions
The body of evidence accumulated to date suggests that generative AI can achieve substitution-level accuracy most reliably in well-defined, domain-specific assessment tasks — particularly short-answer scoring, rubric-aligned essay evaluation, and automated feedback generation in STEM and language learning contexts — where efficiency gains relative to human grading are consistently documented and reproducible. Evidence is considerably weaker, however, in areas that matter most for equitable and consequential assessment: fairness across demographically and linguistically diverse student populations, transparency of scoring rationale, readiness for high-stakes deployment, and the capacity to augment educational quality beyond mere reductions in grader workload. A persistent and underappreciated tension runs through this literature, namely that aggregate accuracy metrics can appear reassuring while concealing systematic differential performance across student subgroups, and that fluency-sensitive models remain vulnerable to construct-irrelevant features such as surface sophistication, verbosity, or stylistic conformity — risks that adequate mean-level benchmarks are structurally ill-equipped to detect. Against this backdrop, three priorities should orient the next generation of research: first, fairness and bias evaluation must be established as a non-negotiable, routine component of any new system’s validation rather than an occasional supplementary analysis; second, principled human-in-the-loop frameworks are needed that specify when, how, and under what conditions human judgment should override or complement AI scoring, thereby preserving efficiency benefits without inducing the automation bias that uncritical deference to algorithmic outputs tends to produce; third, and most fundamentally, the field must shift its evaluative centre of gravity away from accuracy benchmarks assessed in isolation and toward longitudinal evidence that links the quality of AI-generated evaluation directly to measurable improvements in student learning outcomes. Until these gaps are addressed systematically, the promise of generative AI in educational assessment will remain partially fulfilled — technically impressive in narrow conditions, but insufficiently validated for the broader, higher-stakes contexts in which its adoption is rapidly accelerating.
review: LLMs were used to support grading and feedback processes in institutional workflows, with lecturers describing a tension between perceived efficiency gains and the need for additional human judgment [209].