AI for Education: A Living Catalog of Emerging Practice

Author

Kim Uittenhove

Published

July 28, 2026

What is the Living Catalog?

The Living Catalog of Generative AI for Education is an evolving resource designed to support both researchers and educators in exploring how AI is being used in teaching and learning. Developed by the Center for Learning Sciences at EPFL, the catalog maps a range of applications of generative AI in education.

Our goal is to point users to relevant scientific literature and examples of real-world practice. By doing so, we aim to provide both a starting point and practical guidance for those seeking to implement generative AI in pedagogically sound ways.

The ISAR Framework

The application of generative AI in teaching and learning is characterized by new affordances (i.e., automatisation, adaptation, personalization, interactivity) as well as new risks (i.e., applied without oversight, GenAI can harm learning and teaching). This balance between new affordances and new risks can be understood through the ISAR model (Bauer et al., 2025).

Substitution AI replaces existing tasks or tools with functionally equivalent results.
Augmentation AI adds functional value through pedagogical alignment or adaptive support.
Redefinition AI enables entirely new educational experiences, previously unattainable, by leveraging interactivity and simulation capabilities.
Inversion — a risk at every level: uncritical adoption can undermine the very educational goals AI is meant to support.

How the Catalog is Built

The catalog is produced through a pipeline spanning systematic literature retrieval, LLM-based relevance screening, full-text annotation, analysis, and AI-assisted synthesis. Each stage is built around explicit checkpoints for human input: researchers define which practices to track, how papers should be read and annotated, what should be analyzed, and how the synthesis should be structured.

A distinctive feature is the annotation and analysis approach. The pipeline uses a researcher-defined taxonomy to annotate each excerpt with a section and rhetorical functions, which are subsequently mapped to the evidence requirements for the analysis fields. This makes evidence retrieval tractable and creates a direct avenue for continuous improvement, as domain experts can refine the annotation taxonomy and evidence requirements for analysis fields.

Click on any stage below to learn more.

Expert-Defined Practices & Human-in-the-Loop

  • Practice Definition: Identifying and defining the key emerging practices in generative AI for education, grounded in research, theory, and real-world experience.
  • Annotation taxonomy: Experts define how scientific papers should be annotated.
  • Analysis fields: Experts define the specific analysis questions and evidence requirements for each field.
  • Synthesis structure: Experts provide a synthesis structure template.
  • Decision checkpoints: At key stages, a human reviews the outputs and can modify any part of the extraction process to get the desired results.

Evidence from the Field

  • Open-Access Literature: The catalog is built exclusively on open-access scientific literature.
  • Database Search: For each practice, tailored queries retrieve references from three major academic databases — Web of Science, Scopus, and Semantic Scholar.
  • Relevance Screening: An LLM reviews each article's abstract to assess whether the envisioned practice is a primary focus of the study. Only relevant articles advance to the next stage.
  • Full-Text Retrieval: For screened-in articles, the pipeline collects the full text from publishers or open-access repositories.

Annotation, Analysis & Synthesis

  • Rhetorical annotation: Each article's full text is split into consecutive windows. With the full set of windows as context, the LLM classifies each window by structural section and rhetorical function (e.g., describes participants, reports results, acknowledges limitations, etc.) — a content-agnostic map of how the paper is organized.
  • Field-level analysis: For each expert-specified analysis field, a defined set of rhetorical functions determines which windows are relevant candidate evidence. The LLM then answers that field's specific question using the matching windows, and reports exactly which excerpts it relied on — producing one structured, traceable answer per field per paper.
  • ISAR classification: The interpretation of the artifact presented in each paper is categorized according to the ISAR model.
  • Synthesis generation: The per-field answers are interpreted and structured according to an expert-provided synthesis template, wherein each point in the narrative is linked back to the corresponding paper. A human expert then evaluates and adapts the synthesized content and references before finalization.

Unlike standard retrieval-augmented generation (RAG), where relevance is determined by semantic similarity, this approach puts the expert in control: an expert-defined taxonomy explicitly determines how to annotate papers and the expert decides which rhetorical functions are relevant to which analysis questions, based on domain knowledge of how scientific papers are structured. Evidence retrieval becomes tractable, not dependent on the vagaries of embedding distance.

Automatic Updating Pipeline

The pipeline is designed for automation and traceability, enabling the catalog to stay current as the field evolves. The same queries and procedures are applied consistently across all papers and update cycles, and each processing step produces a timestamped dataset — making it possible to track what changed between runs, compare results across versions, and reproduce any output.

Stage Description
Retrieval Search queries are run against Web of Science, Scopus, and Semantic Scholar; results are merged into an up-to-date reference dataset.
Screening & Full Text Abstracts are screened for relevance by an LLM; full texts are retrieved for screened-in articles.
Annotation Each article is split into windows and classified by structural section and rhetorical function, producing a content-agnostic map of the paper usable across all annotation fields.
Analysis For each analysis field, matching annotated windows are filtered by rhetorical function and the LLM answers the field's question, citing the specific excerpts it used as evidence.
Synthesis Field-level answers are assembled into a narrative synthesis following the expert-provided template, and reviewed by a human expert.

🚨 Human decision checkpoints between stages allow for review before proceeding.


AI for Education

Explore each practice below — select a practice card to read its synthesis of key applications. The full set of underlying papers for each practice can be downloaded as a CSV at the top of its synthesis.

Select a practice
Data verification

Papers cited with a reference number in the synthesis below have been checked by a human expert. However, besides pointing to available literature, we do not evaluate the scientific quality of methodological rigor of studies. Please consult original publications and apply your own critical judgment.

Applications for Assessment Design

Assessment design in the era of AI is characterized by new affordances (i.e., automated, adaptive, personalized, interactive assessments) as well as new risks (i.e., applied without oversight, GenAI can reduce engagement and pedagogical quality). This balance between new affordances and new risks can be understood through the ISAR model (Bauer et al., 2025). Substitution refers to how AI can straightforwardly replace human item authoring with functionally equivalent results. Augmentation captures usage that adds functional value by integrating pedagogical frameworks or adaptive delivery tailored to individual learners. Redefinition describes ways in which we can fully leverage the capabilities of generative AI to enable new assessment formats, such as interactive assessments and simulations. Inversion is a risk across all levels: AI tools adopted uncritically can actively undermine educational goals when assessment quality is insufficient or when there are negative outcomes for learning and engagement. We first map studies by intended innovation level, before turning to beneficial outcomes and potential inversion signals.

Substitution: Generating Assessment Items and Exercises

This section looks at how generative AI can be applied to substitute and scale up assessment design, by enabling extensive and rapid production of examination items that replicate the form and function of human-authored questions. Table 1 provides an overview of studies that leverage generative AI in this way. An important concern for these applications is how the quality of AI-generated items compares to human-authored items. Low assessment quality can undermine evaluation processes and learning outcomes, thereby constituting a risk for inversion, which we extensively address in measures and outcomes.

Table 1. Studies using generative AI for assessment item generation, by discipline.
Field References
Medical and Health Sciences Including structured question formats such as MCQs: [123], [165], [294], [499], [516], [437], [236], [351], [396], [101], [104], [184], [142], [400], [443], [327], [402], [382], [479], [477], [302], [383], [332], [381], [212], [232], [103], [436], [355], [196], [87], [168], [218], [262], [121], [216], [328], [352], [337], [322] Including clinical cases or vignettes: [270], [264], [111], [88], [154], [128], [82], [83], [296], [494], [169], [345], [408], [377], [389], [259], [498]
Computer Science and Engineering [129], [197], [489], [242], [126], [164], [97], [98], [241], [179], [223]
Language and Literacy [119], [162], [273], [217], [203], [125], [195], [472]
Mathematics and Physics [233], [127], [240], [431]
Domain-General Education across Primary, Secondary, and Tertiary Settings [99], [331], [493], [419], [140], [251]

Augmentation: Enhanced and Adaptive Assessment Design

This section focuses on how generative AI can be used to add functional value to assessment design. Three main augmentation mechanisms emerge from the literature.

First, AI-generated assessment items can be aligned with cognitive frameworks like Bloom’s taxonomy to ensure systematic alignment with pedagogical intentions. For example, BloomLLM was fine-tuned on expert-authored questions labeled with Bloom levels, enabling the model to generate assessment items at specified cognitive depths [181]. TwinStar uses a dual-agent architecture in which a separate cognitive-level evaluation agent iteratively feeds back Bloom-alignment scores to a generator agent [274].

Second, adaptive delivery tailors item selection to individual learner characteristics in real time. For example, AdaptiveGPT adjusts question difficulty in real time based on evaluated student responses and presents individualised learning trajectories to teachers [175]. Another study tested a mathematics exercise generator that aligned question difficulty with student cognitive profiles [404].

Third, item generation can be integrated within learning management systems, intelligent tutoring systems, and automated assessment pipelines, affording more accessible and integrated learning opportunities. For example, the Learnix platform integrates GPT-4 to generate MCQs and provide real-time feedback on coding problems within a unified eLearning environment [157].

Table 2. Studies using augmentation mechanisms in AI-assisted assessment design.
Augmentation Mechanism References
Pedagogical Alignment [181], [274], [249], [333], [132], [144], [73], [178], [506], [126], [376], [245], [272], [100], [114], [72], [248], [177], [434], [81], [344], [146], [141], [93], [384], [295], [478], [189], [474], [194], [108], [106]
Adaptive Delivery [175], [263], [126], [404], [440], [372], [300], [326], [344], [304], [120]
Integration [157] [145], [330], [225], [453], [32], [201], [433]

Redefinition: Interactivity and Simulation

At the Redefinition level, generative AI enables new assessment formats that fully leverage the interactivity and simulation afforded by generative AI. These assessments can involve elements like scenarios, branching based on the choices of the examinee, and interactive examination, where the examiner’s next question depends on the candidate’s prior answer. For example, in one study in medicine, decision points are inserted at clinically strategic moments so that the narrative branches in response to student choices [515]. Another study developed an application that generated dynamic clinical case vignettes complete with patient history, vital signs, laboratory values, and imaging findings, and then engaged in an interactive dialogue that guides the examinee through diagnosis, operative planning, and complication management, before delivering structured feedback [122]. In language education, the LEARN system pairs assessment items with a voice-driven adaptive dialogue manager that elicits and evaluates oral responses in multiple languages, creating a multimodal assessment format virtually impossible with pre-AI technology [467]. Another approach links assessment to dynamic content: one driving simulator combines 360-degree spherical video with a visual language model to generate situational test questions anchored to specific moments in the video stream [364].

Table 3. Studies employing redefinition-level AI assessment formats.
Innovation References
Interactive and Simulation [122], [265], [280], [380], [467], [133] [515], [364], [90]

Measures and Outcomes

Whether AI-assisted assessment design is beneficial hinges on the quality of the generated items. As can be seen in Table 4, a considerable subset of studies use expert ratings to evaluate AI-generated items. A smaller set of studies attempt to ascertain whether AI-generated items can meet the psychometric standards of human-authored items, for example by using Classical Test Theory. Very few studies compare student performance on AI-generated and human-authored assessments. Table 4 shows that many studies that assess the quality of AI-generated assessments report inversion risks.

Before turning to inversion risks, we describe some studies with favorable evidence for the quality of AI-generated assessments. One study found that expert median ratings for AI-generated items across five quality domains were statistically identical to human-authored items [443], whereas another study even found that ChatGPT-generated items received higher expert ratings than faculty-prepared items on most quality parameters including clinical appropriateness, distractor reasonableness, and appropriate difficulty [106]. Psychometric studies also frequently reported encouraging results. For example, one study found that AI-generated MCQs presented higher internal reliability, a better-balanced difficulty distribution, and an improved discrimination index compared with a human-generated exam, [100]. However, these results were obtained after human preselection of suitable AI-generated exam items, which is the case for many psychometric studies reporting good results.

At the level of individual item performance, four studies found no significant difference between student scores on AI-generated and human-authored items. For example, one study found no significant difference in exam scores between an AI-generated and faculty-authored assessment with a strong positive correlation between the two [216].

However, four other studies did find performance differences. For example, one study found that AI-generated versions yielded higher mean grades than classic exams but with worrying psychometric values, including a poor difficulty distribution, limited discriminatory power, and inadvertent cues [295], while another study found that AI-generated MCQs were associated with a significant drop in mean midterm scores and greater score variability, which the authors interpreted as evidence of reduced grade inflation and better student differentiation [100].

Table 4. Studies evaluating the quality of AI-generated assessment items, by method.
Method References
Psychometric Evaluation [100], [272], [196], [154], [477], [245], Inversion risks: [383], [328], [82], [317], [216], [264], [337], [189], [81], [262], [474], [295], [126],
Quantitative expert ratings [443], [106], [142], [479], [169], [259], [249], [178], [389], [195], [333], [103], [402], [345], [273], [217], [132], [436], [351], Inversion risks: [377], [332], [396], [382], [302] [146] [81], [218], [119], [493], [262], [88], [474], [295], [352], [232], [141], [376], [127], [241], [73], [121], [179], [111], [203], [101], [408],
Student Performance No difference in performance: [216], [236], [477], [196] Difference in performance: [295], [100], [337], [264]

Discrepancies between student performance on AI-generated and human-authored exams necessitates an evaluation of inversion risk. Alongside evidence of adequacy, a substantial body of studies documents systematic quality deficits in AI-generated items (See Table 5). When left unchecked, these deficits can produce assessments that actively harm learning outcomes — the defining condition of Inversion in the ISAR model.

The most consistently reported deficit concerns low cognitive demand and difficulty: AI-generated items disproportionately target lower-order cognitive skills like recall, and tend to be easier than human-authored items. One study found that 95% of AI-generated questions targeted knowledge recall and only 5% higher-order understanding or analysis [302]. Even when explicitly asked to generate items at specific levels, Apply and Conceptual Knowledge levels showed higher generation precision, suggesting that AI-generation is most viable for those lower levels [146]. A related but distinct issue is difficulty calibration: when AI systems are asked to estimate item difficulty, their ratings show systematic bias compared to empirically derived parameters [300].

Another pervasive issue is low distractor quality. For example, one study found that ChatGPT items had 33.75% nonfunctional distractors versus 13.75% for faculty items, with lower mean distractor effectiveness [262]. Another study found that human-authored distractors received much higher plausibility ratings than LLM-generated ones [127].

Factual inaccuracies, hallucinations and logical inconsistencies also remain an issue. For example, one study found 22% of AI items contained factual or conceptual errors and three items were removed mid-semester after students identified errors [248]. A study of mathematics assessment found frequent calculation errors, incorrect roots, and wrong correct answers in AI-generated items [240]. Some studies also noted an inability for generative AI to produce usable image-based items [437], complicating multimodal assessments.

A final quality deficit reported in the literature is psychometric underperformance. One study found AI items had significantly lower mean discrimination index between high-performing and low-performing students, with 56% of AI items failing acceptable discrimination thresholds versus 26% of human items, and three AI items with negative discrimination [328].

Table 5. Studies documenting quality deficits in AI-generated assessment items, by deficit type.
Quality Deficit References
Low Cognitive Demand and Difficulty [302] [146] [300] [81], [248], [218], [119], [493], [472], [262], [88], [194], [251], [264], [474], [352], [232], [295], [126], [141], [337], [102], [165], [376].
Low Distractor Quality [262], [127], [328], [241], [381], [376], [73], [121], [382], [352], [179], [111], [396], [99], [93], [218], [223], [203], [162], [128], [472], [264], [101], [327], [383], [332]
Factual Inaccuracies, Hallucinations and Logical Inconsistencies [248], [240], [408], [380], [376], [87], [122], [126], [218], [230], [102], [233], [145], [88]; [128], [168], [431], Including inability to produce usable image-based items: [437], [382], [162]
Psychometric Underperformance [328] [82], [317], [216], [262], [264], [337], [189]

Impact on Learning and Engagement

Beyond evaluating the quality of AI-generated assessment items, only few studies have looked at impact on student outcomes such as learning and engagement. Evaluating these outcomes is crucial for understanding inversion risk.

Regarding learning performance, a randomised comparison in three university computing courses found that students who engaged with AI-generated assessments achieved significantly higher final exam scores than those who did not, with positive correlations between number of assessment attempts and exam performance [98]. A study using LLM-generated personalised assessments reported a significantly higher post-test mean versus for controls, with especially pronounced gains for initially lower-performing students [372]. A driving-simulator study reported a 27% improvement in test accuracy after LLM-assisted simulator-based instruction compared with traditional video instruction [364].

Beyond learning, several studies documented positive effects on engagement. The driving-simulator study reported a 32% increase in self-reported engagement and confidence [364]. A language learning platform study reported strong engagement and high enjoyment in students [467]. However, other studies noted negative effects of AI-generated assessments on engagement. For example, one study found that engagement with AI-generated assessments was more superficial (typically completed once) compared with instructor-created content (frequent repeated completions suggesting mastery-oriented use) [330]. A study of AI-generated speaking assessments found that emotional, cognitive, and social engagement were significantly lower with AI than human interlocutors, raising concerns about learning effectiveness in communicative contexts [133]. Lower student engagement with AI-generated assessment constitutes an important inversion signal that should not be ignored, and the contexts where reduced engagement may manifest should be better understood.

Table 6. Studies measuring student impact of AI-generated assessment.
Student Impact References
Learning [98], [372], [364], [114], [82], [164], [429], [144]
Engagement Positive effects: [364], [467], [294], [408], [280], [372] Negative effects: [330], [133]
Student Perception [280], [327], [265]

Recommendations to Improve AI-Enhanced Assessment Design

Practical recommendations to improve the performance of AI-enhanced assessment design have converged around strategies to improve item quality, improve the capability of the assessment design pipeline, and how to design a human-AI workflow (for a detailed overview, see [96]).

Improve item quality

Several strategies allow to improve the quality of AI-generated assessment items, as well as their relevance in relation to course material, alignment with course objectives, while also reducing hallucinations.

Prompt engineering is a consistently reported lever to improve performance. Model outputs vary substantially with prompt wording [121], and novice users who lack prompt-engineering training tend to obtain suboptimal results [101]. Therefore authors across domains treat this as the primary means of improving item quality. Structured prompting and domain-specific prompt templates (e.g., specifying subject, topic, learner level, Bloom’s taxonomy level, distractor constraints) substantially raise the relevance and cognitive alignment of generated items, and reduce common errors such as implausible distractors [97], [104]. Misconception-oriented prompting (i.e., explicitly instructing the model to generate distractors that reflect common student errors) has also been shown to increase the proportion of high-quality distractors in mathematics and computing contexts [242], [241]. Few-shot exemplars are widely recommended as a practical complement to structured prompting. Providing sample questions alongside the prompt biases the model toward desired formats and quality levels [219]. Chain-of-thought (CoT) decomposition (i.e., prompting the model to reason step-by-step through item construction) has furthermore been shown to improve generation quality [127]. Authors also caution that prompt quality is itself a skill requiring professional development: training teachers to craft clear, detailed, goal-oriented prompts is a prerequisite for reliable outputs [86].

Retrieval-augmented generation (RAG) is recommended as a mechanism for reducing hallucination, ensuring curriculum alignment, and incorporating up-to-date knowledge, by enabling the LLM to retrieve relevant passages from course materials or curated knowledge at inference time. Implementations can make use of vector databases [423] or knowledge graphs [110]. Authors recommend careful knowledge-base governance: maintaining master documents in plain text, applying version control, scheduling faculty audits to flag obsolete or missing content, and ensuring source traceability so that every generated item can be traced to an auditable passage in the source material [419], [435]. Creating additional metadata (e.g., knowledge components, key concepts, learning objectives, cognitive complexity ratings, etc.) can further enhance the LLM’s ability to generate appropriate assessment items. For example, one study used syllabus-specific vector databases storing Bloom’s taxonomy classifications and course outcomes to constrain generation [333].

Improve capability of assessment design pipeline

Choice of model is one important factor influencing the performance of generative AI in assessment design (see Table 7). Results obtained with one model snapshot may not replicate with a later version or a different model. The conditions under which adequate quality is most reliably achieved include the use of more capable models (GPT-4 and its successors consistently outperform earlier or smaller models). However, using more capable models is also less sustainable, especially if smaller models can achieve adequate performance.

Multiple agent architecture is recommended by a growing body of work, and consists of decomposing the assessment-generation pipeline into specialised components rather than relying on a single monolithic model call. Different agents could handle tasks related to item generation, pedagogical review, and student error modelling. For example, a dual-agent architecture separating a question-generation model from a cognitive-level evaluation model has been shown to improve Bloom’s-level adherence [274]. More elaborate pipelines assign even more specialised roles (e.g., item generator, distractor critic, difficulty calibrator, and hallucination checker) and use optimisation methods to identify the most cost-effective agent sequences [452]. Indeed, authors consistently note that multi-agent pipelines increase quality but also computational cost, and that minimising agent calls while preserving quality gains is an important practical constraint.

Fine-tuning represent a final and more technically difficult avenue. Fine-tuning allows to repurpose models to develop more specific capabilities, and to perform these capabilities at a lower computational cost. For example, TwinStar uses two lightweight fine-tuned models and attains equal or better question quality while using far fewer parameters and negligible memory compared with large LLMs, making it less costly, more efficiently deployed, and more environmentally friendly [274].

Design Human-AI Workflow

While for low-stakes formative quizzes, and initial item-bank seeding, it may be argued that automated generation and validation of assessment items is adequate, other domains, such as summative examinatons, require a human-in-the-loop workflow. Such a workflow typically involves three checkpoints [477]. First, prompt engineering aligns generation with pedagogical objectives. Second, expert review of generated items ensures that the generated items indeed match pedagogical objectives and contain factually correct information and high-quality distractors. Authors recommend using calibrated rubrics, inter-rater reliability checks, and consensus among multiple reviewers to reduce subjectivity [477]. Third, psychometric validation, including difficulty index, discrimination index, item-total correlation, and IRT analyses, is recommended before deployment; authors caution against circular validation in which the same LLM that generates items also evaluates them [295], [300]. Expert validation and psychometric evaluation should inform an iterative refinement pipeline before items are deployed [154], [214].

Beyond merely integrating human expertise in the process, generative AI applications can also be envisioned as a co-design partner to help humans analyze and improve assessment design [92]. For example, one study designed a generative AI application with a structured workflow inspired by an instructional design framework, operating in several modes, ranging from autonomous to guided to full co-pilot, enabling flexible control over the degree of human involvement [424].

Finally, responsible implementation principles require attention to equity and access, accountability, and the development of AI and assessment literacy across all stakeholders (see [512] for a procedural guidance taxonomy-of validity, reliability, safety, security, accountability, transparency, explainability, privacy, and fairness—for organizations adopting generative AI in assessment contexts). Concerning equity and access, authors recommend performing regular bias audits of generated content, and designing prompts that explicitly incorporate equity considerations [133]. Furthermore, equitable access to AI tools needs to be ensured [488]. Transparency mechanisms are recommended with instructors (e.g., exposing semantic-similarity scores to instructors, providing citations for every assessment item) to prevent black-box assessment design [334], as well as with students (i.e., clarity about AI involvement in assessment design) to maintain trust [212]. Accountability is framed not merely as error-checking but as a professional responsibility: educators must remain the final arbiters of curricular alignment and cognitive-level targeting, and institutions should establish clear accountability frameworks and policies for appropriate use [81]. AI literacy and assessment literacy are identified as mutually reinforcing competencies that must be developed in parallel. Faculty training should cover not only prompt engineering and model selection but also the critical appraisal of AI outputs, interpretation of psychometric indices, and the judgment required to decide when AI-generated items are not fit for deployment [134].

Table 7. Recommendations for AI-enhanced assessment design, with supporting references.
Recommendation References
Choice of Model [135], [245], [196], [296], [165], [72], [83], [178], [142]; [389], [32], [97], [419]; [355], [437], [382], [217], [144]
Prompt Engineering Impact of prompting: [121], [106], [423], [101] Structured prompting: [186], [97], [104], [208], [81], [178], [474], [274], [376] Leveraging misconceptions: [242], [241] Few-shot exemplars: [219], [223], [402], [127] Chain-of-thought: [171], [127], [32], [97], [369] Teacher training: [86], [420], [443], [476]
RAG [423], [110], [188], [364], [419], [274], [80], [333], [249], [111], [435] Metadata enrichment: [333], [95], [137], [367], [358], [135], [258]
Multiple Agents [274], [452], [280], [80], [296], [249], [99], [203], [132], [171]
Fine-tuning [274], [372], [127], [80], [209], [394], [158]
Human-in-the-loop Workflow: [214], [477], [295], [300], [345], [337], [219], [87], [82], [248], [111], [264], [81], [245] Iterative refinement: [154], [196], [498], [383], [452], [429] Co-design: [92], [424], [505], [80]
Responsible Implementation Equity and access: [133], [488], [505], [453], [319] Transparency: [334], [212], [419], [396] Accountability: [81], [212], [134], [85], [334], [419] AI and assessment literacy: [134], [238], [443]

Limitations & Gaps

The main limitation concerns the lack of evidence for pedagogical effectiveness and insitutional impact of AI-assisted assessment design.

Limitations in Sample, Scope, and Generalizability. Most studies rely on small samples in a single discipline. Exposure periods tend to be short. In many cases, deployment in real classrooms is lacking, leaving key questions about pedagogical impact, student engagement, and discriminative power under real examination conditions unanswered. A review of generative AI in K-12 education reinforces this finding, noting a lack of concrete experiments and longitudinal validation [243].

Limitations in Evaluation of Quality and Impact. To assure assessment quality, expert review of AI-generated items is the dominant approach (see Table 4). However, inter-rater reliability is repeatedly flagged as problematic [232], [443], [302], [380], [245], [408], [189]. Moreover. measurement of impact of AI-assisted assessment design on student outcomes, such as learning and engagement, is limited, with a small set of studies reporting mostly positive findings. Long-term retention is almost universally unexamined: short assessment windows preclude conclusions about whether AI-supported assessment produces durable learning. The absence of learner-centered validation — collecting student feedback on the utility, clarity, or perceived fairness of AI-generated items — is also noted in several studies leaving a significant gap between expert-rated quality and the student experience of assessment.

Limitations in Generative AI Technology. Studies documented extensive evidence of several quality deficits in AI-generated assessment items (see Table 5). The manifestation of these issues depends on several factors, including model version and prompt engineering, but also point to structural limitations of current generative approaches. Several mitigation strategies have shown promise in the literature, and are thoroughly discussed in the next section.

Conclusion and Future Directions

The field of generative AI in assessment design has advanced rapidly from early proof-of-concept demonstrations of automated item authoring toward more ambitious efforts to augment the quality of assessment instruments and, in some cases, redefine what assessment formats can look like . Yet this momentum has not been without cost: empirical evidence now makes visible a pattern of Inversion risks, wherein the uncritical adoption of generative tools produces quality failures, including factual inaccuracies, lower cognitive levels, problematic distractor quality, and psychometric underperformance. Evidence on student outcomes compounds this picture: while some AI-assisted applications show learning gains, engagement can be lower with AI-based assessments compared to human interlocutors.

There are several priorities for future research. First, item quality must be systematically improved through advances in prompt engineering, retrieval-augmented generation techniques that ground model outputs in verified domain knowledge, and multi-agent architectures. Computational sustainability also deserves greater attention in this regard: architectures combining smaller finetuned models with RAG pipelines and rich metadata can approach the quality of large-model systems at substantially lower cost, making capable AI-assisted assessment design viable outside well-resourced institutions.

Second, the field must move decisively beyond its current preoccupation with multiple-choice formats to develop generative approaches for multimodal, adaptive, and simulation-based assessments that can capture complex competencies and better reflect authentic performance contexts.

Third, the current literature on the impact of AI-generated assessment on student learning and engagement remains sparse, short-term, and largely confined to multiple-choice formats. Building a rigorous longitudinal evidence base across contexts is a prerequisite for informed scaling.

Lastly, professional capacity needs to be developed. As AI tools take on a larger share of assessment authoring, the ability of educators to critically evaluate, refine, and take responsibility for AI-generated items cannot be assumed — it must be actively cultivated through sustained investment in AI literacy, assessment literacy, and clear institutional accountability frameworks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

How to Cite This Catalog

If you use this catalog in your research or work, please use the citation information below.


Uittenhove, K. (2025). AI for Education: A Living Catalog of Emerging Practice (Version v1.0.0). Zenodo. https://doi.org/10.5281/zenodo.15633017

Uittenhove, Kim. AI for Education: A Living Catalog of Emerging Practice. Version v1.0.0, Zenodo, 10 June 2025, doi:10.5281/zenodo.15633017.

Uittenhove, Kim. 2025. "AI for Education: A Living Catalog of Emerging Practice." Zenodo. https://doi.org/10.5281/zenodo.15633017.

Uittenhove, K. (2025) 'AI for Education: A Living Catalog of Emerging Practice', Zenodo. doi: 10.5281/zenodo.15633017.

1. Uittenhove K. AI for Education: A Living Catalog of Emerging Practice [Internet]. Zenodo; 2025 Jun 10. Available from: https://doi.org/10.5281/zenodo.15633017

[1] K. Uittenhove, "AI for Education: A Living Catalog of Emerging Practice," Zenodo, Jun. 10, 2025. doi: 10.5281/zenodo.15633017.
@misc{uittenhove_2025,
  author       = {Uittenhove, Kim},
  title        = {{AI for Education: A Living Catalog of Emerging Practice}},
  month        = {June},
  year         = {2025},
  publisher    = {Zenodo},
  version      = {v1.0.0},
  doi          = {10.5281/zenodo.15633017},
  url          = {[https://doi.org/10.5281/zenodo.15633017](https://doi.org/10.5281/zenodo.15633017)}
}