Phase 0-2: Schema cleanup, typed relations, event-driven automation

- Phase 0: AGENTS.md cleanup (dedup quotes, renumber sections, merge qmd)
- Phase 1: typed relations (manage-relations.py, graph-search.py, check-staleness.py, detect-conflicts.py)
- Phase 2: frontmatter validator, weekly lint, knowledge promotion, git hooks
- Fix .gitignore to track tools/ and .githooks/
- Fix git remote URL (remove plaintext token)
- New wiki pages: 504 pages, 34 raw sources
This commit is contained in:
hehaiguang1123
2026-07-01 08:05:43 +08:00
parent e544d6e04a
commit a6f05ab2d5
1067 changed files with 522992 additions and 819 deletions
+18
View File
@@ -0,0 +1,18 @@
{
"page_1": " Contents lists available at ScienceDirect\nComputers and Education: Artificial Intelligence\njournal homepage: www.sciencedirect.com/journal/computers-and-education-artificial-intelligence \nLarge language models in education: a systematic review of empirical \napplications, benefits, and challenges\nYuhong Shi iD, Kun Yu, Yifei Dong, Fang Chen\nData Science Institute, Faculty of Engineering and Information Technology, University of Technology Sydney, Ultimo, NSW 2007, Australia\nH I G H L I G H T S\n• Reviews 88 empirical studies on LLM applications in education, selected from 3344 publications since ChatGPTs release (Nov 2022Mar 2025).\n• Identifies six key LLM applications, with Intelligent Tutoring Systems being the most common.\n• Empirical evidence shows that LLMs enhance academic performance, engagement, and cognitive abilities.\n• Identifies critical concerns: over-reliance, fairness, privacy, and technical issues.\n• Mixed findings on cognitive development demand longitudinal research studies.\nA R T I C L E I N F O\nKeywords:\nLarge language models \nEducational technology \nArtificial intelligence \nChatGPT \nEmpirical studiesA B S T R A C T\nThe rapid advancement of Large Language Models (LLMs), particularly following the release of ChatGPT in \nNovember 2022, has significantly transformed educational methodologies. This systematic review aims to syn­ \nthesize empirical studies published between November 2022 and March 2025, examining the implementation \nand effectiveness of LLMs in educational settings. 88 empirical studies identified key applications, benefits, and \nchallenges associated with LLM integration in education. Our findings reveal that LLMs are utilized across various \neducational contexts in six primary applications, with Intelligent Tutoring Systems being particularly prominent. \nThe benefits include improved academic performance, increased student engagement, enhanced accessibility, \noptimized resource utilization, and strengthened cognitive and skill development. However, challenges such as \nstudent over-reliance on AI, technical reliability issues, assessment fairness, and privacy concerns were identi­ \nfied. This review provides educators, researchers, and policymakers with evidence-based insights and practical \nguidance for effective LLM integration, contributing to the ongoing transformation of teaching and learning in \nthe era of Generative Artificial Intelligence (GenAI) technology.\n1 . Introduction\nLarge language models (LLMs) are Artificial Intelligence (AI) sys­ \ntems designed to process and generate human-like text by learning \npatterns from extensive training data ( Xu, Chen & Miao , 2024 ). The \nwidespread adoption of sophisticated LLMs, particularly following the \nrelease of ChatGPT in November 2022, has initiated a new era of educa­ \ntion enhanced by Generative Artificial Intelligence (GenAI), prompting \nextensive research into their educational applications and implications \n(Zarei et al. , 2024 ). T",
"page_2": "Y. Shi, K. Yu, Y. Dong et al.\ntheir progress, and reflect on their understanding through personalized \nfeedback and metacognitive prompts ( Fan et al. , 2025 ). Concurrently, \nLLMs align with both Cognitive Load Theory ( Sweller , 1988 ) and the \nZone of Proximal Development (ZPD) ( Vygotsky , 1978 ) through their \nadaptive capabilities. Specifically, Vygotsky s ZPD theory conceptual­ \nizes the gap between what learners can accomplish independently and \nwhat they can achieve with guidance, while Sweller s Cognitive Load \nTheory posits that learning effectiveness depends on how instructional \ndesign manages the limited capacity of working memory by balancing \nintrinsic, extraneous, and germane cognitive load. LLM-integrated sys­ \ntems adjust response complexity and break down intricate concepts into \nmanageable components while simultaneously assessing learners cur­ \nrent understanding to provide appropriately challenging content that \nneither overwhelms nor underwhelms students, thereby supporting ef­ \nfective conceptual understanding ( Yang et al. , 2024 ; Yunianto et al. , \n2024 ).\nWhile previous literature reviews have explored LLM applications in \neducation ( Samala et al. , 2024 ; Zarei et al. , 2024 ), the systematic syn­ \nthesis of empirical implementations in actual classroom environments \nremains underexplored. Recent developments in LLM-based educational \ntools have generated valuable empirical evidence about practical appli­ \ncations and outcomes, however, these studies remain scattered across \ndifferent educational applications, necessitating a systematic synthesis \nof implementation approaches. This review synthesizes emerging em­ \npirical findings across six functional categories of LLM applications, \nincluding Chatbots, Learning Content Generation, Automated Assessment \nand Feedback, Task Support Tools, Learning Support Tools, and Intelligent \nTutoring Systems , providing evidence-based insights for implementation \nacross diverse educational settings.\nIn particular, this systematic review adhered to the Preferred \nReporting Items for Systematic Reviews and Meta-Analyses (PRISMA) \nguidelines ( Liberati et al. , 2009 ), spanning from the launch of ChatGPT \nto when this review was conducted. This time frame captures the trans­ \nformative period following ChatGPTs release. The release of ChatGPT \nmarked a significant milestone in educational technology, catalyzing \nunprecedented development and implementation of LLMs, including \nadvanced architectures such as GPT-J, BLOOM, GODEL, and Cohere \nSandbox ( Lin et al. , 2024 ). This period witnessed significant shifts in \npedagogical approaches and technological integration, characterized by \nan exponential increase in empirical research examining the effective­ \nness, limitations, and implications of LLM-based tools ( Chen , 2023 ; Lyu \net al. , 2024 ; Yuan et al. , 2023 ). Focusing on this period, our review \nencompasses emerging research findings and practical application",
"page_3": "Y. Shi, K. Yu, Y. Dong et al.\nTable 1 \nSummary of related systematic review studies on LLMs in education.\nCitation Domain Coverage period Contributions\nChatGPT in English Language Teaching (ELT) ( Adipat , \n2025 ) ELT 20202024 ChatGPTs opportunities, challenges, and ethical \nconsiderations.\nLLM in Higher Education ( Chhina et al. , 2023 ) HE 20182023 Benefits and challenges of LLMs in higher \neducation. \nChatGPT in ELT ( Wang, Hanafi Zaid, et al. , 2024 ) ELT 20232024 Opportunities, challenges, and trends in applying \nChatGPT in ELT. \nOpen-Source LLMs in Education ( Lin et al. , 2024 ) General Education 20232024 Open-source LLMs and their suitability for ed­\nucational applications in English-speaking \ncontexts. \nLLMs in Medical Education ( Lucas et al. , 2024 ) Medical Education 20222024 LLMs application, opportunities and challenges in \nmedical education. \nChatGPTs Pros and Cons in Learning and Teaching \n(Samala et al. , 2024 ) General Education 20182023 Advantages and disadvantages of ChatGPT and its \nrole as a supportive learning tool. \nChatGPT in Healthcare Education ( Sallam , 2023 ) Healthcare \nEducation20222023 Benefits and limitations of ChatGPTs utility in \nhealthcare education, scientific research, and \npractice.\nChatGPT on Critical Thinking (Melisa et al., 2025 ) HE 20232024 ChatGPTs impact on students critical thinking \nand evaluative judgment. \nLLM-based Code Generation Models ( Cambaz & Zhang , \n2024 )Programming \nEducation20182023 LLM-based code generation models in teaching \nand learning practices, along with their character­ \nistics, evaluation indicators, and considerations for \nintegration. \nAI and LLMs in Learning and Teaching ( Xu, Gu & Lu , 2024 ) HE 20202024 Benefits and challenges of AI and LLMs in \neducational practice.\n• RQ2: What empirical evidence exists regarding the benefits and pos­\nitive impacts of LLM integration on teaching and learning outcomes \nin educational environments?\n• RQ3: What are the key challenges and concerns associated with \nimplementing LLM-based educational technologies, as identified \nthrough empirical research?\nTo systematically address these research questions, studies were clas­ \nsified according to three analytical dimensions following established \nsystematic review practices ( Imran & Almusharraf , 2023 ; Lo, Hew & \nJong, 2024 ; Lo, Yu, et al. , 2024 ). The classification framework en­ \ncompassed geographical distribution to identify regional publication \npatterns, domain classification to capture discipline-specific applica­ \ntions, and educational context categorization to identify level-specific \npedagogical applications.\nA systematic coding procedure was implemented with one researcher \nserving as the primary coder for all included studies, while a three -\nmember research team engaged in regular collaborative verification \nsessions to ensure validity and interpretive consistency through consen­ \nsus adjudication. Each study wa",
"page_4": "Y. Shi, K. Yu, Y. Dong et al.\n• Empirical studies presenting original, evidence-based findings from \nthe systematic collection and analysis of data involving real par­ \nticipants, such as students or educators interacting with LLM-based \neducational systems in authentic or controlled settings.\n• Publications within the specified time frame, from November 2022 \nto March 2025.\n• Studies conducted in formal educational contexts, including K12 \nand higher education institutions.\n• Research focusing on LLM applications in education.\n• Studies that employ systematic evaluation methods using quantita­\ntive, qualitative, or mixed-method approaches and report evidence \nsuch as participants feedback, learning outcomes, behavioral data, \nor expert evaluations.\n3.2.2 . Exclusion criteria\n• Studies presenting non-empirical content, including theoretical \nframeworks, conceptual papers, literature reviews, opinion pieces, \nposition papers, and meta-analyses, were excluded to maintain focus \non primary findings.\n• Research conducted in non-traditional educational contexts, such as \nprofessional training, informal learning settings, or lifelong learn­ \ning initiatives, was excluded to ensure consistency in educational \nenvironment analysis.\n• Publications that failed to demonstrate methodological rigor through \nclear methodology, empirical data, or explicit educational impli­ \ncations were excluded to maintain the quality standards of the \nreview.\n• Studies focusing solely on machine learning, deep learning, or tradi­\ntional AI applications without incorporating LLM components were \nexcluded to maintain a specific focus on contemporary LLM-based \neducational technologies.4 . Result\nThe systematic review process was executed in multiple phases \nfollowing the PRISMA guidelines. The initial search yielded 3344 po­ \ntentially relevant publications across the selected databases. Following \nthe predetermined inclusion criteria emphasizing peer-reviewed publi­ \ncations, a preliminary screening was conducted to remove duplicates \nand non-journal or non-conference literature, resulting in 2022 unique \njournal or conference papers for further evaluation. The subsequent \nscreening phase systematically assessed titles and abstracts against the \npredefined research questions and inclusion criteria. This process led to a \nfurther exclusion of 214 records that did not align with the reviews focus \non educational settings, yielding 1808 publications for full-text examina­ \ntion. In the final screening phase, a rigorous full-text analysis evaluated \nthe methodological robustness and empirical validity of the remaining \nstudies. This comprehensive assessment resulted in the exclusion of 1720 \npublications that either lacked empirical research methodology or did \nnot focus on LLM applications in education. The final corpus comprised \n88 studies that demonstrated robust empirical evidence for LLM ap­ \nplications in educational contexts. The complete screen",
"page_5": "Y. Shi, K. Yu, Y. Dong et al.\nFig. 2. Distribution of studies by region.\nUSAs contributions (n = 16), with Canada adding one study. The \nEuropean region demonstrates greater geographical diversity, contribut­ \ning 18.2 % (n = 16) of the studies, distributed across 12 countries. \nGermany leads the European contributions with five studies, followed \nby Switzerland (n = 2), and nine other countries contributing one study \neach. Additional contributions come from Africa (4.6 %, n = 4), Oceania \n(2.3 %, n = 2), and South America (1.1 %, n = 1). This distribution \nhighlights a significant concentration of research output in a few key \ncountries, particularly China and the USA, which in total account for \n43.2 % of all studies.\n4.2 . Domain distribution\nFig. 3 shows the distribution of studies by domain, revealing that \nthe Computer and Technology domain dominates the field with 38.6 % \n(n = 34) of the total studies. Within this domain, Programming repre­ \nsents the largest subset with 19 studies, followed by Computer Science (n \n= 6), while Computational Thinking and Data Science each contribute \ntwo studies, and Cybersecurity, Software Engineering, AI, VR, and Data \nLiteracy each contribute one study. The Language Learning and Writing \ndomain emerges as the second most prominent domain, accounting for \n25.0 % (n = 22) of the studies, with Writing (n = 16) and non-writing \nLanguage Learning (n = 6) comprising this category. The STEM domain \nconstitutes 18.2 % (n = 16) of the studies. Mathematics dominates this \ncategory with eight studies, followed by general STEM education (n = \n3), Engineering and Science (n = 2 each), and Biology (n = 1). The \nOthers category comprises 9.1 % (n = 8) of the studies and includes di­ \nverse domains such as Healthcare (n = 3), Reading (n = 2), and single \nstudies in Art, Management, and Storytelling Skills. Additionally, stud­ \nies without specified educational domains account for another 9.1 % (n = 8) of the total. This distribution highlights a significant concen­ \ntration of research in Computer and Technology and Language Learning \nand Writing, which collectively represent 63.6 % of the total studies.\n4.3 . Educational contexts distribution\nThe educational level distribution of the included studies focused \nexclusively on formal education settings, with Higher Education com­ \nprising 76.1 % (n = 67) and K-12 representing 23.9 % (n = 21) of \nthe studies. We excluded studies conducted outside formal K-12 and \nhigher education institutions, including vocational training and lifelong \nlearning, as well as informal learning settings. This approach aligns with \nour primary research question, examining empirical implementations of \nLLMs in established educational settings. By concentrating on formal \nlearning environments, where educational goals and assessment meth­ \nods are consistently applied, we can make more reliable generalizations \nabout LLM educational applications.\n4.4 . Summary of LLM-based application",
"page_6": "Y. Shi, K. Yu, Y. Dong et al.\nTable 3 \nCompilation of research studies proposing LLM-based applications.\nApplications References\nChatbots Chen and Chen (2023 ); Chen (2023 ); Chen, Juan, et al. (2024 ); Looi and Jia (2025 ); Mohammed \net al. (2025 ) \nLearning Content Generation Bezirhan and von Davier (2023 ); Chen et al. (2023 ); Choi et al. (2024 ); del Carpio Gutierrez et al. \n(2024 ); Elmourabit et al. (2024 ); Liu et al. (2025 ); Logacheva et al. (2024 ); Norberg et al. (2024 ); \nPesovski et al. (2024 ) \nAutomated Assessment and \nFeedbackAhmed et al. (2025 ); Alshammari (2025 ); Bergerhoff et al. (2025 ); Cagliero et al. (2024 ); Choi and \nKim (2025 ); Dai et al. (2023 ); Hadyaoui and Cheniti-Belcadhi (2024 ); Hutt et al. (2024 ); Hwang and \nNurtantyana (2022 ); Jansen et al. (2025 ); Meyer et al. (2024 ); Nguyen and Park (2025 ); Ouyang \net al. (2024 ); Riazi and Rooshenas (2025 ); Singh et al. (2024 ); Su et al. (2024 ); Wang et al. (2025 ); \nXiao and Liu (2025 ) \nTask Support Tools Fan et al. (2025 ); Hou et al. (2024 ); Li (2023 ); Ohm et al. (2024 ); Ouaazki et al. (2024 ); Qureshi \n(2023 ); Torres (2023 ); Tsao et al. (2024 ); Vishnumolakala et al. (2024 ); Xiao and Liu (2025 ); Yang \net al. (2024 ); Yunianto et al. (2024 ); Zhang et al. (2023 ); Zhu et al. (2025 ) \nLearning Support Tools Alvarez (2024 ); Bešlić et al. (2024 ); Canonigo (2024 ); Chen, Jiang, et al. (2024 ); Feng and Wang \n(2025 ); Gao et al. (2024 ); Gasaymeh and Almohtadi (2024 ); Hong et al. (2024 ); Jin et al. (2024 ); \nKumar et al. (2024 ); Liffiton et al. (2024 ); Mi and Li (2025 ); Oktarin et al. (2024 ); Pears et al. (2024 ); \nPeng et al. (2023 ); Qin et al. (2024 ); Tang et al. (2024 ); Xu and Liu (2025 ); Yuan et al. (2023 ); Zhou \net al. (2024 ) \nIntelligent Tutoring Systems Abolnejadian et al. (2024 ); Baba et al. (2024 ); Chun et al. (2025 ); Civit et al. (2024 ); Faruqui et al. \n(2024 ); Lai and Lin (2025 ); Liu et al. (2024 ); Lyu et al. (2024 ); Mejia-Domenzain et al. (2025 ); Nam \net al. (2024 ); Nutalapati et al. (2024 ); Panwale and Vijayakumar (2025 ); Park et al. (2024 ); Pian \net al. (2024 ); Santhosh et al. (2024 ); Sarshartehrani et al. (2024 ); Schmucker et al. (2024 ); Son et al. \n(2024 ); Soudi et al. (2023 ); Teng et al. (2024 ); Wei and Yan (2024 ); Wong et al. (2023 )\nenhance the learning experience by providing explanations and supple­ \nmentary materials related to knowledge acquisition. Intelligent Tutoring \nSystems integrate LLM capabilities to create adaptive learning environ­ \nments that respond to individual student needs, learning styles, and \nlearning paces. Together, these applications represent a comprehensive \necosystem of LLM-powered tools that impact student engagement and \nlearning outcomes across educational settings. While these categories \nprovide a ",
"page_7": "Y. Shi, K. Yu, Y. Dong et al.\npose significant barriers to educational implementation. Additionally, \nconcerns about over-reliance on auto-generated content indicate peda­ \ngogical challenges related to maintaining student autonomy and critical \nthinking skills. The relatively limited discussion of assessment evolu­ \ntion and fairness (n = 8) and privacy and security (n = 6) suggests that \nthese areas may require more research attention, particularly given their \nimportance for institutional adoption and ethical deployment. These \nchallenges collectively underscore the need for careful consideration of \ntechnical, ethical, and pedagogical factors when implementing LLMs in \neducational settings.\n5 . Discussion\nLLMs possess distinctive technological features that fundamentally \nreshape educational possibilities through their natural language process­ \ning capabilities for interactive dialogue engagement, generative abilities \nfor dynamic content creation, and adaptive and immediate responsive­ \nness to individual learning needs. These capabilities directly translate \ninto diverse educational applications, including interactive chatbots, \ncontent generation tools, automated assessment systems, task and learn­ \ning support tools, and intelligent tutoring systems. These applications \nfind robust theoretical grounding in established educational frame­ \nworks. Constructivist learning theory ( Vygotsky , 1978 ) validates active \nknowledge creation, while Vygotsky s ZPD enables personalized scaf­ \nfolding across multiple contexts. Cognitive Load Theory ( Sweller , 1988 ) \nguides information presentation and processing, and SRL ( Zimmerman , \n2000 ) informs autonomous skill development. Active Learning princi­ \nples ( Bonwell & Eison , 1991 ; Prince , 2004 ) promote engaged participa­ \ntion, and Hattie and Timperley s (2007 ) Feedback Model directs effective \nfeedback mechanisms.\nThese theoretical alignments translate into educational benefits, \nincluding enhanced cognitive development, improved academic perfor­ \nmance, and increased student motivation and engagement. However, \nthe widespread adoption of LLMs in educational contexts also raises \nconcerns regarding over-reliance on LLM-generated responses, techni­ \ncal reliability of their output quality, assessment fairness, and privacy \nissues related to student data collection and usage.\n5.1 . Applications\nWhile previous reviews have examined LLM applications in educa­ \ntion ( Cambaz & Zhang , 2024 ; Lucas et al. , 2024 ; Samala et al. , 2024 ), \nour categorization provides a comprehensive six-category analytical \nframework ( Table A2 outlines the details) that captures technological \ncapabilities and pedagogical applications across multiple educational \ndomains while exclusively focusing on empirical evidence from actual \nimplementations, rather than being limited to specific disciplines or \ncombining theoretical and empirical studies. These categories represent \nthe broad spectrum of ",
"page_8": "Y. Shi, K. Yu, Y. Dong et al.\nerrors remain ( Cagliero et al. , 2024 ). Beyond assessment, LLM feedback \nrooted in constructivist principles demonstrates positive cognitive and \naffective-motivational outcomes through adaptive, in-depth guidance \n(Alshammari , 2025 ; Meyer et al. , 2024 ; Wang et al. , 2025 ), providing \nwriting revision suggestions and examples ( Hwang and Nurtantyana , \n2022 ; Xiao and Liu , 2025 ), code-specific guidance that fosters au­ \ntonomous motivation ( Choi and Kim , 2025 ; Ouyang et al. , 2024 ), \nand iterative refinement that progressively addresses issues ( Riazi and \nRooshenas , 2025 ).\nHowever, LLM-based systems cannot fully replace human pedagog­ \nical relationships ( Ahmed et al. , 2025 ). Evidence reveals engagement \nchallenges, as Jansen et al. (2025 ) found that approximately half of stu­ \ndents make no revisions after receiving ChatGPT-generated feedback, \nand research shows that GPTs scaffolding quality varies considerably \ndepending on problem complexity ( Singh et al. , 2024 ). These patterns \nunderscore that pedagogical effectiveness depends not only on feedback \nquality but also on developing students capacities to critically evaluate \nand appropriately implement LLM-generated feedback ( Su et al. , 2024 ).\n5.1.4 . Task support tools\nLLMs function effectively as task support tools across diverse educa­ \ntional domains, with their scaffolding capabilities broadly characterized \nacross pre-task preparation, task implementation, and review dimen­ \nsions. In pre-task preparation, these systems assist with brainstorming, \nideation, planning and outlining, step-by-step goal recommendations, \nand rapid draft prototyping ( Li, 2023 ; Tsao et al. , 2024 ; Vishnumolakala \net al. , 2024 ; Xiao and Liu , 2025 ; Zhang et al. , 2023 ). During task \nimplementation, LLMs provide real-time guidance with detailed expla­ \nnations, pseudocode guidance, progressive hints for different levels of \nsupport, and personalized content ( Hou et al. , 2024 ; Ouaazki et al. , \n2024 ; Qureshi , 2023 ; Torres , 2023 ; Yang et al. , 2024 ; Yunianto et al. , \n2024 ; Zhu et al. , 2025 ). Review-oriented support includes revision while \npreserving students distinctive styles and original ideas, grammar cor­ \nrection, code debugging, and manuscript refinement ( Fan et al. , 2025 ; \nLi, 2023 ; Vishnumolakala et al. , 2024 ; Xiao and Liu , 2025 ; Yunianto \net al. , 2024 ).\nWhile LLMs demonstrate effectiveness in immediate task support \nand completion, considerations arise regarding their impact on deeper \nlearning outcomes and knowledge transfer capabilities ( Fan et al. , 2025 ; \nOhm et al. , 2024 ). This does not suggest limiting LLM use in academic \neducation, but rather employing it with caution. Recommendations in­ \nclude implementing permissive policies that encourage transparent LLM \nuse (Ohm et al. , 2024 ), helping learners develop self-regulated learning \nskills and maintain metacognitive activity ( F",
"page_9": "Y. Shi, K. Yu, Y. Dong et al.\n5.2 . Benefits\nThe growing body of empirical research on LLM integration in \neducation reveals substantial benefits and significant uncertainties. \nThese technologies demonstrate improved student performance, with \nenhanced test scores and faster task completion across language learn­ \ning, programming, and mathematics ( Feng and Wang , 2025 ; Torres , \n2023 ; Zhu et al. , 2025 ). LLMs offer key educational benefits, including \nenhanced motivation through personalized learning, expanded accessi­ \nbility for remote learners, efficient resource optimization via streamlined \nassessment and content creation ( Bezirhan and von Davier , 2023 ), and \npotential for developing critical thinking and self-regulated learning \ncapabilities ( Chen et al. , 2023 ).\nWhile these academic performance benefits are notable, they can \nbe compromised depending on usage approaches. Students who accept \nAI-generated responses without reflection may experience diminished \ncritical thinking ( Qin et al. , 2024 ), whereas those who critically eval­ \nuate AI-generated suggestions and actively verify ChatGPTs output \ndemonstrate enhanced cognitive development ( Hadyaoui and Cheniti -\nBelcadhi , 2024 ). Consequently, emerging research emphasizes the need \nfor pedagogically grounded integration strategies, which are essential \nfor developing evidence-based guidelines that maximize benefits while \nmitigating risks to authentic learning.\n5.2.1 . Academic performance\nWhile previous reviews ( Adipat , 2025 ; Lucas et al. , 2024 ) have ex­ \namined LLM applications in isolated educational domains, our review \nintegrates quantitative performance metrics across language learning, \nwriting, and programming contexts, revealing convergent patterns of \npersonalized learning effectiveness, as evidenced in Table A1 .\nIn language learning, LLMs provide real-time feedback on pronun­ \nciation, facilitate human-computer dialogues for oral practice, generate \ncustomized reading materials, offer writing corrections, deliver trans­ \nlation assistance, and create professional communication simulations. \nFeng and Wang (2025 )s semester-long study documented a 20 % im­ \nprovement in Chinese college students English proficiency compared \nto a 1.2 % improvement in control groups applying traditional ap­ \nproaches. Writing instruction exhibits similar advantages as students \ndemonstrate enhanced syntactic complexity, accuracy, and overall qual­ \nity compared to the control group that received traditional writing \ninstruction from teachers ( Li, 2023 ; Oktarin et al. , 2024 ; Wong et al. , \n2023 ; Xiao and Liu , 2025 ; Zhou et al. , 2024 ). Similar impressive out­ \ncomes were observed in programming education, with ChatGPT-assisted \nlearners achieving reduced completion times, higher success rates, and \nimproved mean scores across multiple studies ( Choi and Kim , 2025 ; \nGasaymeh and Almohtadi , 2024 ; Lyu et al. , 2024 ; Nutalapati et al. , 2024 ; ",
"page_10": "Y. Shi, K. Yu, Y. Dong et al.\nefficient management of high-volume tasks such as programming feed­ \nback ( Torres , 2023 ) and cost-effective development of learning resources \n(Pears et al. , 2024 ). This optimization capability extends to instruc­ \ntional design, where LLMs can address student questions and doubts \n(Teng et al. , 2024 ), while also enabling teachers to create pedagogically \nmeaningful dialogues from existing lectures ( Choi et al. , 2024 ). By au­ \ntomating time-intensive tasks, LLMs allow educators to redirect efforts \ntoward deeper, more meaningful student interactions ( Abolnejadian \net al. , 2024 ), thereby enhancing the overall quality of education while \nmaintaining cost-effectiveness.\n5.3 . Challenges and concerns\nThe architectural design of LLMs shapes both their capabilities and \nlimitations. LLMs, such as ChatGPT, utilize transformer architectures \nand deep neural networks to process vast amounts of text data and \nlearn language patterns ( Tayan et al. , 2024 ). While this design enables \nsophisticated tasks like language generation and contextual reasoning, \nthe “black box” nature of their billions of parameters raises concerns \nabout transparency and reliability. These LLMs reliance on pre-trained \ndata can lead to hallucinations, generating plausible but incorrect re­ \nsponses when facing ambiguous or novel contexts. Furthermore, biases \npresent in training data may be perpetuated through the models re­ \nsponses, raising assessment fairness concerns ( Elmourabit et al. , 2024 ). \nThese technical limitations become particularly problematic when stu­ \ndents uncritically accept LLM outputs, fostering over-reliance ( Cagliero \net al., 2024 ). Finally, the collection and analysis of student data for per­ \nsonalized learning present privacy and security risks that require robust \nprotective measures. The following analysis synthesizes empirical find­ \nings across four critical dimensions, including over-reliance, technical \nreliability, assessment fairness, and privacy considerations.\n5.3.1 . Over-reliance\nThe integration of LLMs in educational environments presents sev­ \neral critical challenges, with students potential overdependence on \nautomated feedback emerging as a primary concern ( Cagliero et al. , \n2024 ). This overdependence manifests in various problematic behaviors, \nincluding students accepting LLM-generated responses without ques­ \ntioning or critical evaluation ( Mi and Li , 2025 ) and engaging in excessive \nuse of these tools ( Lai and Lin , 2025 ). The issue becomes particu­ \nlarly problematic when these LLMs provide helpful responses even to \npoorly articulated queries, inadvertently reinforcing students unsophis­ \nticated communication behaviors. Moreover, the immediate availability \nof assistance may discourage the development of crucial debugging and \nanalytical skills, especially in specialized domains such as data science \neducation ( Yuan et al. , 2023 ). Indeed, research evidence ",
"page_11": "Y. Shi, K. Yu, Y. Dong et al.\neducational domains. They foster motivation and engagement through \npersonalized tutoring and adaptive feedback that cultivate enthusi­ \nasm and proactive learning. They also promote cognitive development \nby facilitating critical thinking and problem-solving skills through \nscaffolding, reflective reasoning, and interactive learning experiences. \nAdditionally, LLMs improve accessibility by transcending geographi­ \ncal and temporal constraints through 24/7 availability. Finally, they \noptimize resources by automating routine tasks such as assessment \nand content creation, freeing educators to focus on deeper student \nengagement and meaningful pedagogical interactions.\n5.4.2 . Concerns in LLM applications\nDespite their potential, the review identifies critical challenges that \ndemand careful consideration. The primary concern involves poten­ \ntial student over-reliance, as the very convenience and responsiveness \nthat make LLMs engaging may inadvertently compromise students de­ \nvelopment of independent problem-solving and analytical skills. This \npedagogical risk becomes particularly problematic when coupled with \nsignificant technical limitations, such as inconsistent accuracy and \nhallucinations, which can mislead students, especially in high-stakes \neducational contexts where precision is essential. These technical uncer­ \ntainties compound automated assessment challenges, as LLM-generated \nevaluations may introduce subtle biases and inequitable treatment that \nundermine the fairness and transparency essential to educational assess­ \nment. Notably, privacy and ethical considerations underscore the urgent \nneed for comprehensive, secure, and accountable frameworks governing \nLLM implementation in educational settings.\n6 . Conclusion\nThis systematic review identified six key applications of LLMs, with \nIntelligent Tutoring Systems emerging as particularly prominent. The \nfindings reveal both benefits and concerns in LLM implementation. The \nmultifaceted benefits highlight LLMs potential to enhance academic \nperformance, increase student motivation, improve accessibility, and \noptimize resource utilization. However, the review also underscores crit­ \nical concerns, including student over-reliance, technical unreliability, \nfairness in assessment, and privacy risks.\nLLMs hold immense potential to revolutionize educational environ­ \nments, as demonstrated by the steadily growing number of publications \neach month since the launch of ChatGPT, but their optimal use re­ \nquires thoughtful contextualization and collaboration between systems \nand human educators. Research consistently indicates that LLMs should \nsupplement, rather than replace, traditional teaching methods to achieve \nthe best outcomes ( Bešlić et al. , 2024 ; Soudi et al. , 2023 ). These models \nare particularly effective in scenarios that benefit from scalability and \npersonalization, such as automated feedback, adaptive learning paths, \nand l",
"page_12": "Y. Shi, K. Yu, Y. Dong et al. \nAppendix \nSee Tables A1 and A2 \nTable A1 \nSummary of studies reporting meta-analytic metrics on LLM benefits in education. \nCitation Statistical analysis Educational data analytics \nReported \nMetric Effect Size Value Statistical \nSignificance Domain Application Performance Outcomes \nAlshammari (2025 ) T-test Cohens 𝑑 = 1.412 𝑡(52) = 5 .19, 𝑝 < 0.001 Programming ChatGPT-enhanced Adaptive \nE-learning System Programming test scores \nAlvarez (2024 ) T-test Not reported 𝑡 = 14 .453, 𝑝 < 0.001 Math LLM-powered tutor Math tests scores \nBaba et al. (2024 ) T-test Not reported 𝑡 = 5.246.02 (four \nsubjects), 𝑝 < 0.001 Multi LLM-powered personalized learning system Knowledge test scores \nCanonigo (2024 ) T-test Cohens 𝑑 = 2.36 𝑡(60) = 6 .673, 𝑝 < \n0.05, CI = [1.13, \n2.10] Math GeoGebra and ChatGPT Conceptual un­derstanding test scores \nChen (2023 ) ANCOVA Not reported 𝐹 = 5.94, 𝑝 = 0.003 Science GPT-3.5-Turbo Conceptual test scores \nChen, Juan, et al. \n(2024 ) ANCOVA 𝑓 = 0.148 𝐹 = 12 .140, 𝑝 = 0.001 Language ChatGPT-powered Learning Tool Proficiency test scores \nChen, Jiang, et al. \n(2024 ) T-test Not reported 𝑝 < 0.001 Bio-Inspired Design \n(BID) LLMs-driven tool Quiz scores \nChoi and Kim \n(2025 ) ANCOVA 𝜂2 = 0.062 𝑝 < 0.001 Programming LLM-based programming \nlearning environment Programming ability test scores \nChun et al. (2025 ) T-test Cohens 𝑑 = 0.86 𝑝 = 0.879 Health LLM-powered digital \ntextbooks Exam scores \nFan et al. (2025 ) ANOVA Not reported CI = [3.858, \n0.083], 𝑝 = 0.037 Writing GPT 4.0 Essay scores \nFeng and Wang \n(2025 ) T-test Not reported 𝑝 < 0.001 Language ChatGPT Exam scores \nGasaymeh and \nAlmohtadi (2024 ) T-test Not reported 𝑡(72) = 2 .063, 𝑝 = \n0.04 Programming ChatGPT Skills test scores \nHwang and \nNurtantyana (2022 ) ANCOVA 𝜂2 = 0.157 𝐹 (1, 68) = 12 .37, \n𝑝 < 0.01 Writing GPT-2- powered app Essay scores \nLi (2023 ) T-test Not reported 𝑡(41) = 2 .2.502, 𝑝 < \n0.05 Writing ChatGPT Writing test score \nLiu et al. (2024 ) T-test Not reported 𝑡(29) = 12 .5(𝑡𝑜𝑡𝑎𝑙 ), \n𝑝 < 0.05 Language ChatGPT-powered ITS English skills test scores \nLooi and Jia (2025 ) T-test Not reported 𝑡(50) = 9 .220, 𝑝 < \n0.001, 95 % CI = \n[2.030, 1.304] Writing ChatGPT Summary assessment \nscores \nLyu et al. (2024 ) T-test Not reported 𝑡 = 2 .847, 𝑝 = 0.009 Programming LLM-powered assistant Exam scores \nMeyer et al. (2024 ) Regression Analysis 𝑑 = 0.19 𝑝 = 0.042 Writing GPT-3.5-Turbo Essay revision scores \nMi and Li (2025 ) T-test Not reported 𝑡(6) = 4 .889, 𝑝 = 0.003 Not specified SparkDesk Project scores \nMohammed et al. \n(2025 ) ANOVA 𝜂2 = 0.859 𝐹 (3, 201) = 408 .793, \n𝑝 < 0.001 Computer Science ChatGPT Computer Education Achievement Test \n(CEAT) \nNutalapati et al. \n(2024 ) T-test Cohens 𝑑 = 1.04 𝑡(998) = 16 .42, 𝑝 < \n0.001 Programming Fine-tuned GPT-3.5 Coding assessment scores \nOktarin et al. \n(2024 ) T-test Not reported 𝑡(48) = 6 .028, 𝑝 < \n0.001 Writing ChatGPT English writing test \nscores \nPanwale and ",
"page_13": "Y. Shi, K. Yu, Y. Dong et al.\nTable A2 \nSummary of reviewed studies grouped by application type.\nApplication type Domains LLM types Outcome measures Citations (n)\nChatbots Science, Language Learning, \nWriting, Computer ScienceGPT-3.5-Turbo, ChatGPT Performance, motivation, \nobservations, grit and growth \nmindset scalesChen and Chen (2023 ); Chen (2023 ); Chen, Juan, et al. \n(2024 ); Looi and Jia (2025 ); Mohammed et al. (2025 ) \n(n = 5)\nLearning Content \nGenerationReading, Storytelling Skills, \nSTEM, Programming, \nComputer Science, \nMathematics, Software \nEngineeringGPT-3, GPT-4, GPT-3.5 -\nTurbo, ChatGPT, GPT-4 \nTurboContent quality, performance, \nusage patterns, learning \nexperienceBezirhan and von Davier (2023 ); Chen et al. (2023 ); \nChoi et al. (2024 ); del Carpio Gutierrez et al. (2024 ); \nElmourabit et al. (2024 ); Liu et al. (2025 ); Logacheva \net al. (2024 ); Norberg et al. (2024 ); Pesovski et al. \n(2024 ) (n = 9)\nAutomated \nAssessment and \nFeedbackProgramming, Data Science, \nMathematics, Writing, \nScience, Computer Science, \nAIGPT-4 Turbo, ChatGPT, GPT -\n4, GPT-2, GPT-3.5 Turbo, \nClaude 3.5 Sonnet, Gemini \n1.5 Flash, GPT-4o, GPT-3.5Performance, grading \naccuracy, engagement, per­ \nceptions, learning experience, \nmotivation, cognitive skills, \nlearning behavior, feedback \nqualityAhmed et al. (2025 ); Alshammari (2025 ); Bergerhoff \net al. (2025 ); Cagliero et al. (2024 ); Choi and Kim \n(2025 ); Dai et al. (2023 ); Hadyaoui and Cheniti -\nBelcadhi (2024 ); Hutt et al. (2024 ); Hwang and \nNurtantyana (2022 ); Jansen et al. (2025 ); Meyer et al. \n(2024 ); Nguyen and Park (2025 ); Ouyang et al. (2024 ); \nRiazi and Rooshenas (2025 ); Singh et al. (2024 ); Su \net al. (2024 ); Wang et al. (2025 ); Xiao and Liu (2025 ) \n(n = 18)\nTask Support \nToolsWriting, Programming, \nCybersecurity, \nComputational Thinking, \nComputer Science, \nEngineering, MathematicsGPT-4, ChatGPT, GPT -\n3, GPT-3.5 Turbo, \nRAG-powered toolsPerformance, motivation, \nSRL process metrics, en­ \ngagement, usage patterns, \nlearning experienceFan et al. (2025 ); Hou et al. (2024 ); Li (2023 ); Ohm \net al. (2024 ); Ouaazki et al. (2024 ); Qureshi (2023 ); \nTorres (2023 ); Tsao et al. (2024 ); Vishnumolakala \net al. (2024 ); Xiao and Liu (2025 ); Yang et al. (2024 ); \nYunianto et al. (2024 ); Zhang et al. (2023 ); Zhu et al. \n(2025 ) (n = 14)\nLearning Support \nToolsMathematics, Engineering, \nLanguage Learning, Data \nLiteracy, Programming, \nComputational Thinking, \nComputer Science, Writing, \nHealthcare, Management, \nReadingChatGPT, GPT-4, GPT-3.5 \nTurbo, GPT-3, Sparkdesk, \nT5, MixQGPerformance, conceptual un­ \nderstanding, engagement, \nperceptions, writing and \nsoft skills, cognitive skills, \nlearning experience, usage \npatterns, observationsAlvarez (2024 ); Bešlić et al. (2024 ); Canonigo (2024 ); \nChen, Jiang, et al. (2024 ); Feng and Wang (2025 ); Gao \net al. (2024 ); Gasaymeh an",
"page_14": "Y. Shi, K. Yu, Y. Dong et al. \nChoi, S., & Kim, H. (2025). The impact of a large language model-based programming \nlearning environment on students motivation and programming ability. Education and \nInformation Technologies , 30(6), 81098138. \nChoi, S., Lee, H., Lee, Y., & Kim, J. (2024). Vivid: Human-AI collaborative authoring \nof vicarious dialogues from lecture videos. In Proceedings of the 2024 CHI conference \non human factors in computing systems CHI 24 . New York, NY, USA: Association for \nComputing Machinery. \nChun, J., Kim, J., Kim, H., Lee, G., Cho, S., Kim, C., Chung, Y., & Heo, S. (2025). A compar­\native analysis of on-device AI-driven, self-regulated learning and traditional pedagogy \nin university health sciences education. Applied Sciences (Switzerland) , 15(4). \nCivit, M., Escalona, M. J., Cuadrado, F., & Reyes-de-Cozar, S. (2024). Class integration of \nChatgpt and learning analytics for higher education. Expert Systems , 41(12). \nDai, W., Lin, J., Jin, H., Li, T., Tsai, Y.-S., Gašević, D., & Chen, G. (2023). Can large lan­\nguage models provide feedback to students? A case study on ChatGPT. In 2023 IEEE \ninternational conference on advanced learning technologies (ICALT) (pp. 323325). \ndel Carpio Gutierrez, A., Denny, P., & Luxton-Reilly, A. (2024). Automating personalized \nParsons problems with customized contexts and concepts. In Proceedings of the 2024 on \ninnovation and technology in computer science education V. 1 ITiCSE 2024 (pp. 688694). \nNew York, NY, USA: Association for Computing Machinery. \nDuckworth, A. L., Peterson, C., Matthews, M. D., & Kelly, D. R. (2007). Grit: Perseverance \nand passion for long-term goals. Journal of Personality and Social Psychology , 92(6), \n1087. \nDweck, C. S., Walton, G. M., & Cohen, G. L. (2014). Academic tenacity: Mindsets and skills \nthat promote long-term learning. ERIC Number: ED576649; 43 pp. \nElmourabit, Z., Retbi, A., & El Faddouli, N.-E. (2024). The impact of generative artificial in­\ntelligence on education: A comparative study. In Proceedings of the European conference \non E-learning, ECEL (Vol. 23, pp. 470476). \nFan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). \nBeware of metacognitive laziness: Effects of generative artificial intelligence on learn­ \ning motivation, processes, and performance. British Journal of Educational Technology , \n56(2), 489530. \nFaruqui, S. H. A., Tasnim, N., Basith, I. I., Obeidat, S. M., & Yildiz, F. (2024). Board 46: \nIntegrating AI in higher-education protocol for a pilot study with samcares an adaptive \nlearning hub. In 2024 ASEE annual conference & exposition . \nFeng, Y., & Wang, X. (2025). Exploring the development of Chinese college students \nproficiency in English through chatgpt: An experimental study. In Proceedings of the \n2024 16th international conference on education technology and computers ICETC 24 (pp. \n148154). New York, NY, USA: Association for Computing Mac",
"page_15": "Y. Shi, K. Yu, Y. Dong et al. \nNguyen, H., & Park, S. (2025). Providing automated feedback on formative science as­\nsessments: Uses of multimodal large language models. In Proceedings of the 15th \ninternational learning analytics and knowledge conference LAK 25 (pp. 803809). New \nYork, NY, USA: Association for Computing Machinery. \nNguyen, T. N., & Truong, H. T. (2025). Trends and emerging themes in the effects of \ngenerative artificial intelligence in education: A systematic review. Eurasia Journal of \nMathematics, Science and Technology Education , 21(4), 111. \nNorberg, K. A., Almoubayyed, H., De Ley, L., Murphy, A., Weldon, K., & Ritter, S. (2024). \nRewriting content with GPT-4 to support emerging readers in adaptive mathematics \nsoftware. International Journal of Artificial Intelligence in Education . \nNutalapati, H., Velmurugan, S., & Tiglao, N. M. (2024). Coding buddy: An adaptive \nAI-powered platform for personalized learning. In 2024 international symposium on \nnetworks, computers and communications (ISNCC) (pp. 16). \nOhm, M., Bungartz, C., Boes, F., & Meier, M. (2024). Assessing the impact of large lan­\nguage models on cybersecurity education: A study of chatgpts influence on student \nperformance. In Proceedings of the 19th international conference on availability, reliabil­ \nity and security ARES 24 (pp. 17). New York, NY, USA: Association for Computing \nMachinery. \nOktarin, I. B., Saputri, M. E. E., Magdalena, B., Hastomo, T., & Maximilian, A. (2024). \nLeveraging Chatgpt to enhance students writing skills, engagement, and feedback \nliteracy. Edelweiss Applied Science and Technology , 8(4), 23062319. \nOlugbade, D., Edwards, B. I., & Ojo, O. A. (2024). Facilitating cognitive load management \nand improved learning outcomes and attitudes in middle school technology and voca­ \ntional education through AI chatbot. Journal of Technical Education and Training , 16(3), \n114131. \nOuaazki, A., Bergram, K., Farah, J. C., Gillet, D., & Holzer, A. (2024). Generative AI -\nenabled conversational interaction to support self-directed learning experiences in \ntransversal computational thinking. In Proceedings of the 6th ACM conference on con­ \nversational user interfaces CUI 24 . New York, NY, USA: Association for Computing \nMachinery. \nOuyang, F., Guo, M., Zhang, N., Bai, X., & Jiao, P. (2024). Comparing the effects of instruc­\ntor manual feedback and Chatgpt intelligent feedback on collaborative programming in \nchinas higher education. IEEE Transactions on Learning Technologies , 17, 21732185. \nPanwale, S. B., & Vijayakumar, S. (2025). Evaluating AI-personalized learning interven­\ntions in distance education. International Review of Research in Open and Distributed \nLearning , 26(1), 157174. \nPark, M., Kim, S., Lee, S., Kwon, S., & Kim, K. (2024). Empowering personalized learn­\ning through a conversation-based tutoring system with student modeling. In Extended \nabstracts of the CHI Conference on human factors in comput",
"page_16": "Y. Shi, K. Yu, Y. Dong et al. \nXu, Q., Gu, J., & Lu, J. (2024a). Leveraging artificial intelligence and large language \nmodels for enhanced teaching and learning: A systematic literature review. In 2024 \n13th international conference on computer technologies and development (TechDev) (pp. \n7377). \nXu, X., Chen, Y., & Miao, J. (2024b). Opportunities, challenges, and future directions of \nlarge language models, including Chatgpt in medical education: A systematic scoping \nreview. Journal of Educational Evaluation for Health Professions , 21, 6. \nYang, A. C. M., Lin, J.-Y., Lin, C.-Y., & Ogata, H. (2024). Enhancing Python learning \nwith pytutor: Efficacy of a chatgpt-based intelligent tutoring system in programming \neducation. Computers and Education: Artificial Intelligence , 7. \nYuan, K., Lin, H., Cao, S., Peng, Z., Guo, Q., & Ma, X. (2023). Critrainer: An adaptive \ntraining tool for critical paper reading. In Proceedings of the 36th annual ACM symposium \non user interface software and technology UIST 23 . New York, NY, USA: Association for \nComputing Machinery. \nYunianto, W., Lavicza, Z., Kastner-Hauler, O., & Houghton, T. (2024). Investigating the \nuse of Chatgpt to solve a geogebra based mathematics+computational thinking task \nin a geometry topic. Journal on Mathematics Education , 15(3), 10271052.Zarei, M., Zarei, M., Hamzehzadeh, S., Oliyaei, S., & Hosseini, M.-S. (2025). Chatgpt, \na friend or a foe in medical education: A review of strengths, challenges, and \nopportunities. Shiraz E-Medical Journal [In Press]. \nZhang, Z., Gao, J., Dhaliwal, R. S., & Li, T.-J.-J. (2023). VISAR: A human-AI argumentative \nwriting assistant with visual programming and rapid draft prototyping. In Proceedings \nof the 36th annual ACM symposium on user interface software and technology UIST 23 \n(pp. 130). New York, NY, USA: Association for Computing Machinery. \nZhou, Y., Xu, K., Yin, B., & Liu, N. (2024). Research on the application of digital humans in \nEnglish oral teaching based on AI models. In Proceedings of the 2024 9th international \nconference on distance education and learning ICDEL 24 (pp. 4956). New York, NY, \nUSA: Association for Computing Machinery. \nZhu, W., Xing, W., Lyu, B., Li, C., Zhang, F., & Li, H. (2025). Bridging the gender gap: The \nrole of AI-powered math story creation in learning outcomes. In Proceedings of the 15th \ninternational learning analytics and knowledge conference LAK 25 (pp. 918923). New \nYork, NY, USA: Association for Computing Machinery. \nZimmerman, B. J. (2000). Attaining self-regulation: A social cognitive perspective. In \nHandbook of self-regulation (pp. 1339). Elsevier.Computers and Education: Artiϧcial Intelligence 10 (2026) 100529 \n16 "
}