Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles and JavaScript.
Advertisement
Humanities and Social Sciences Communications volume 13, Article number: 684 (2026)
24k
3
19
Metrics details
With the rapid advancement of generative artificial intelligence (GenAI) technology, the potential educational applications of Chat generative pre-trained transformers (ChatGPT) have attracted significant attention. However, the research on the specific effects of ChatGPT on student learning outcomes and its moderating factors remains insufficient. This study aimed to quantify the effects of ChatGPT on student learning outcomes and explore relevant moderating variables using a meta-analysis approach. The analysis included 35 studies published between 2022 and 2024, involving 4193 participants. The results indicated a moderately positive effect of ChatGPT on student learning outcomes (g = 0.670), significantly enhancing both cognitive and non-cognitive skills. In the analysis of moderating variables, the subject, experimental duration, and instructional mode had significant positive effects on student learning outcomes, whereas educational level and knowledge type did not show significant effects. Additionally, the publication bias test revealed no significant publication bias. This meta-analysis confirmed the effectiveness of ChatGPT in improving student learning outcomes and highlighted the roles of the subjects, experimental duration, and instructional mode as key moderating factors. Despite the risks of sample selection bias and limitations in fully covering the multidimensional moderating factors and higher-order thinking, the findings provided important empirical support for applying ChatGPT in education. Future research should investigate the mechanisms behind the impact of ChatGPT and consider a wider range of moderating factors to thoroughly evaluate its long-term effects on student learning outcomes.
Rapid advancements in artificial intelligence (AI) have significantly transformed the education sector in the current digital era (Yusuf et al., 2024). Generative artificial intelligence (GenAI), a technology capable of creating new content and solutions, has attracted widespread interest due to its potential educational applications. Chat generative pre-trained transformer (ChatGPT) offers personalised learning experiences and enhances teaching methods and learning resources through the simulation, prediction, and automation of complex tasks. Therefore, understanding the impact of ChatGPT on student learning outcomes is essential for optimising its application in education and maximising its benefits. Currently, research on ChatGPT and student learning outcomes mainly concentrates on four areas: the performance of ChatGPT across various academic disciplines (Newton and Xiromeriti, 2023; Susnjak and McIntosh, 2024; Wu et al. 2024); the influence of ChatGPT on education (Kim and Lee, 2023; Pellas, 2023; Urban et al. 2024); teaching, learning, management, and assessment supported by ChatGPT (Murgia et al. 2023; Lee et al. 2024; Mogavi et al. 2023; Luo et al. 2024); and potential issues emerging from the use of ChatGPT (Abramski et al. 2023; Harrison et al. 2023; Othman, 2023). The application of ChatGPT in educational research is gradually gaining importance, with related studies affirming its significance for student learning outcomes. However, the deployment of the ChatGPT application also encounters a series of challenges in research, including result heterogeneity, limited comprehensive evidence, differences in research quality and methodology, and uncertainties in practical applications.
In the field of language learning, Darmawansah et al. (2025) reported a significant improvement in language learning with ChatGPT (effect size η2 = 0.33), but Nakavachara et al. (2024) found that 34% of participants showed no improvement in non-English writing tasks after using ChatGPT. In STEM education, Zhu et al. (2023) reported no difference in participation and interdisciplinary learning between ChatGPT and non-ChatGPT conditions, while Vasconcelos and Santos’ (2023) research highlights ChatGPT’s ability to help learners develop reflective and critical thinking, creativity, problem-solving skills, and conceptual understanding in STEM education. In natural science, Chen and Chang (2024) showed that ChatGPT-assisted game-based learning boosts motivation, reduces cognitive load, and encourages effective science learning. However, the research results of Kosar et al. (2024) indicated that student performance, practical assignments, and mid-term exam scores are not affected by the use of ChatGPT. In social science learning, Liu and Wang (2024) showed that ChatGPT could effectively improve students’ critical thinking skills in English literature classes. Zhou and Kim (2024) found that the experimental group improved in music learning after using the ChatGPT-4 system (ES = 0.097), compared to the control group. While many studies have examined ChatGPT’s effectiveness, the optimal usage conditions remain unclear.
To address these challenges, this study employed a meta-analysis approach to summarise the effects of ChatGPT on students’ learning outcomes based on 35 experimental and quasi-experimental studies. The objectives were to address the heterogeneity of research outcomes, improve the standards for evaluating research quality, provide comprehensive evidence, reduce uncertainty in practical applications, and offer a scientific basis for policymaking. Using this method, we revealed the extent of ChatGPT’s effects across different subjects, educational levels, experimental duration, instructional modes, and types of knowledge. Furthermore, we identified key factors and potential mechanisms influencing these effects, thereby providing clearer and more robust evidence to support the use of ChatGPT in the education sector.
By late 2022, ChatGPT’s emergence brought attention to generative AI and large language models, sparking global interest in academic, media, and public discourse. Its ability to generate data and content offers innovative tools for education, helping teachers personalise instruction and giving students immediate feedback and adaptive pathways to improve learning. However, its effect on learning outcomes remains debated. This study explores ChatGPT’s impact, reviewing studies on its effects and factors influencing student outcomes. The review is divided into two parts: In the first part, we assessed the effects of ChatGPT on student learning outcomes from both qualitative and quantitative research perspectives. In the second part, we investigated the influencing factors across subjects, experimental duration, educational levels, types of knowledge, and instructional modes. The systematic review and meta-analysis of the extant literature provide comprehensive references for educational practitioners and researchers and inform future research and applications.
Qualitative research has been instrumental in understanding the effects of ChatGPT on student learning outcomes. It is suitable for exploratory research in fields lacking theoretical foundations, and can provide a theoretical framework for subsequent quantitative research. This section reviews existing qualitative studies to investigate how the educational use of ChatGPT affects students’ learning experiences, motivation, attitudes, and behaviours.
Studies show that ChatGPT improves students’ learning by offering personalised content and immediate feedback, based on interviews and observations. For example, Rehman et al. (2024) analysed qualitative data and found that ChatGPT offered students a personalised learning experience, evoking their feelings to increase their academic engagement, which positively influenced their perceptions and behaviour towards academic buoyancy. Students reported that the learning process became more interesting and challenging when using ChatGPT-based learning platforms. Albdrani and Al-Shargabi (2023) conducted a qualitative study to investigate the educational effectiveness of ChatGPT in the context of technology and personalised learning. This study demonstrated the effectiveness of ChatGPT in promoting student engagement, learning outcomes, and personalised learning experiences. Students indicated that ChatGPT could provide personalised learning paths based on their progress and needs, thereby increasing their engagement in learning.
Furthermore, qualitative research elucidated the impact of ChatGPT on students’ learning motivations. Through interviews and focus groups, researchers found that ChatGPT boosts students’ interest and intrinsic motivation. Lee et al.(2022) showed that students were more enthusiastic and proactive with AI chatbots. They appreciated the immediate feedback and personalised suggestions, which supported and motivated them academically. Lee et al. (2024) created a ChatGPT-enabled tool integrated into blended learning platforms and courses, significantly raising student interest and motivation. Compared to traditional methods, students using ChatGPT showed greater engagement. Limo et al. (2023) investigated the use of ChatGPT in online tutoring and found that AI-generated instant feedback and personalised suggestions effectively improved student learning outcomes and motivation. Babalola et al. (2024) found a strong positive correlation between students’ interest in critical thinking and their scores in organic chemistry under ChatGPT conditions. Li et al. (2024) used a mixed-methods design to study how post-secondary learners use ChatGPT for self-directed writing. Results showed students’ motivations ranged from curiosity about technology to academic needs. When learners saw ChatGPT’s potential benefits, their initial motivation shifted to task motivation.
Finally, the qualitative research explored the impact of ChatGPT on students’ learning attitudes and behaviours. Through observations and interviews, the researchers found that ChatGPT can promote positive learning attitudes and improve learning behaviours. Rudolph et al. (2023) suggested that students using ChatGPT were more actively involved in classroom discussions and showed a greater willingness to collaborate with their peers on tasks. Furthermore, students generally believed that ChatGPT supported a deeper understanding and mastery of learning content, enhancing their learning confidence. Collecting feedback from teachers and students was essential for this qualitative study. Through interviews and questionnaire surveys, the researchers gained insights into the practical applications and effectiveness of ChatGPT in teaching. Teachers generally recognised that ChatGPT could reduce their workload, enabling them to concentrate on instructional design and student guidance. Students reported that personalised learning experiences and immediate feedback provided by ChatGPT fostered a sense of autonomy and accomplishment during their learning journeys.
In summary, qualitative research shows that ChatGPT profoundly impacts student learning. While it enriches experiences, boosts motivation, and improves attitudes, studies highlight challenges like dependency, privacy issues, and acceptance by teachers and students (Bouzar et al., 2024; Gupta et al., 2023; Strzelecki, 2023).
ChatGPT has transformed teaching by providing automated content, personalised tutoring, and real-time feedback, making it essential in modern educational technology. While research shows ChatGPT can improve students’ learning, its actual impact needs validation through systematic studies. This review examines existing quantitative research—experiments, surveys, and data analyses—focusing on ChatGPT’s role in enhancing academic performance, knowledge, learning speed, efficiency, and motivation. This review provides a scientific foundation and practical guidance for future educational technology and instructional design applications.
Experimental studies are one of the primary methods used to evaluate the impact of ChatGPT on student learning outcomes. Kim and Lee (2023) conducted a trial in which they divided middle school English students into experimental and control groups. The experimental group received ChatGPT-assisted instruction; the control group used traditional methods. Results showed the experimental groups scored 19% higher on finals, indicating ChatGPT’s significant impact on academic performance. Pellas (2023) found that university students using a generative AI platform for digital storytelling significantly outperformed a control group in knowledge, with a 16% in narrative intelligence. Wu et al. (2024) revealed that, in an integrated STEM course for prospective teachers, students guided by a ChatGPT-based learning aid exhibited higher task completion, lower cognitive load, and better critical thinking abilities. Nakavachara et al. (2024) reported that ChatGPT users scored higher and completed tasks faster. However, some studies show that 34% of participants showed no improvement in writing analysis tasks, and 42% in mathematics and data analysis. Dasari et al. (2024) found that students relying only on ChatGPT performed worse in university mathematics education than those with traditional or ChatGPT guidance. These findings suggest ChatGPT’s educational use may need to be combined with traditional methods.
Questionnaire surveys are another commonly used evaluation method for gathering students’ subjective perceptions and feedback on ChatGPT. For instance, Boudouaia et al. (2024) demonstrated that undergraduate students using ChatGPT generally perceived improvements in their learning speed and efficiency. Compared with the pre-test scores, the post-test average scores for perceived usefulness, perceived ease of use, attitude, and behavioural intention showed statistically significant enhancements, indicating that the intervention positively impacted the participants’ views of ChatGPT. Additionally, Chan and Lee (2023) found that Gen Z students in higher education generally held optimistic views about the potential benefits of GenAI, including increased productivity, efficiency, and personalised learning, and expressed their intention to use ChatGPT for educational purposes. Pallivathukal et al. (2024) found that 36.6% of students used ChatGPT for academic purposes; 65.9% believed it significantly reduced the time and effort needed to complete tasks; 47.6% verified the accuracy of the information provided by ChatGPT; and 49% said that they would continue using ChatGPT for academic reasons in the future. Liu et al. (2024) demonstrated that using artificial intelligence for proofreading effectively improved Chinese learners’ English writing and reading skills outside the classroom and boosted their confidence in learning English.
Finally, data analysis is an effective tool for evaluating the impact of ChatGPT on student learning outcomes. Through big data analysis, the effects of ChatGPT in different learning environments and conditions were revealed. Baidoo-Anu and Ansah (2023) found that ChatGPT excels in designing personalised learning paths, creating adaptive platforms that adjust based on student growth and performance. This promotes personalised and successful learning strategies. Students using ChatGPT showed significant improvement in their knowledge mastery. Widianingtyas et al. (2023) explored students’ perceptions of ChatGPT based on the expectancy-value theory. Results showed significant links between students’ ChatGPT knowledge, past use, perceived value, and intention to use, emphasising the role of motivation in adoption decisions. Uddin et al. (2024) discovered that, after using ChatGPT, students’ overall knowledge performance improved by approximately 34%, with written responses becoming more thorough, detailed, and informative. Chan and Zhou (2023) studied how students’ perceptions influence their intention to use ChatGPT in higher education, using expectancy-value theory to assess knowledge, perceived value, and perceived cost. They found a strong positive link between perceived value and intent, and a weak negative link between perceived cost and usage intention.
Nguyen et al. (2024) employed a three-tier learning analytics approach to examine how ten doctoral students interacted with GenAI-assisted tools during academic writing. The findings showed that doctoral students who engaged in iterative and highly interactive processes using GenAI tools generally achieved better results in writing tasks. In contrast, those who used GenAI as a supplementary information source and maintained a linear writing approach often received lower writing scores. Arthur et al. (2024) used the unified theory of acceptance and use of technology 2 to investigate the factors influencing the behavioural intention and usage of ChatGPT among higher education students. The findings indicated that hedonic motivation, performance expectancy, effort expectancy, and social influence were significant predictors of behavioural intention, while behavioural intention and facilitating conditions notably impacted the actual use of ChatGPT. Age and gender acted as moderators in the usage patterns. These findings suggest that ChatGPT has potential in educational applications, necessitating further research on best practices and implementation strategies.
The quantitative evidence confirms ChatGPT’s strong potential to positively influence students’ academic performance, knowledge, learning efficiency, and motivation. However, the qualitative findings provide the crucial context for these outcomes, revealing that the path to achieving them is not straightforward. The very tools that boost motivation and enrich learning experiences (e.g., instant support, personalised interaction) can also lead to over-reliance and superficial processing if not guided properly. Furthermore, the variation in effectiveness across subjects and educational levels, as noted quantitatively, is likely intertwined with the contextual challenges identified qualitatively, including differing levels of teacher and student acceptance and concerns over data privacy (Bouzar et al., 2024; Gupta et al., 2023; Strzelecki, 2023). Therefore, the overall impact of ChatGPT is not merely a function of its use but is profoundly shaped by how it is integrated into the educational environment. Future research should thus focus not only on measuring outcomes in diverse settings but also on developing implementation frameworks that maximise the benefits uncovered by quantitative studies while proactively addressing the risks and challenges illuminated by qualitative enquiry.
As an advanced technology, ChatGPT has increasingly influenced education, transforming teaching and learning. Its impact on Students’ outcomes depends on multiple factors. Understanding these is essential for optimising ChatGPT’s use in education and enhancing results. This article reviews key factors affecting ChatGPT’s effectiveness, including subject matter, duration, educational level, types of knowledge, and instructional modes.
The application of ChatGPT in education yielded varied results across subjects. Kim and Lee (2023) showed that students using ChatGPT in junior high school English improved their test scores by an average of 19%. Boudouaia et al. (2024) found that students in English as a foreign language writing improved in task completion, coherence, cohesion, grammatical and lexical skills, with about a 28% boost in overall writing performance. Wu et al. (2024) reported positive effects of a ChatGPT-supported platform on spoken English, especially in creating authentic environments. Ng et al. (2024) demonstrated that ChatGPT chatbots (SRLbot) enhanced students’ science knowledge, engagement, motivation and reduced anxiety. Hakiki et al. (2023) highlighted that integrating ChatGPT into university technical education significantly improved students’ scores, showing its potential to enhance academic performance. However, Escalante et al. (2023) found that AI-generated feedback did not differ significantly from human tutor feedback in terms of the progress of beginner English learners (ENL). Ju (2023) discovered a decline in students’ learning efficiency with generative AI assistance, with a 25.1% decrease in learning efficiency when completely relying on GenAI for writing tasks, and a 12% decrease when using it for assisted reading. In summary, ChatGPT’s effects vary by subject, with positive impacts on mathematics, science, and literature, but limited improvements in language skills like listening and speaking. While it may be effective in education, further research is needed on best practices and implementation strategies.
The duration of the experiment significantly affected the learning outcomes of ChatGPT. Wu et al. (2024) demonstrated that ChatGPT-based intelligent learning aids could effectively enhance students’ self-regulation and knowledge-construction abilities within a 10-day intervention period. This is particularly true in blended learning environments, where students show greater advantages in stimulating interest, intrinsic motivation, and behavioural engagement. Li and Wang (2023b) found that the human–machine debate method improved students’ performance, attitudes, critical thinking, and problem-solving during a 19-day experiment, but did not show enhancement in innovative thinking from large language models. In a 12-week study, Song and Song (2023) showed that AI-assisted instruction significantly boosted students’ writing skills and motivation compared to controls. Short-term use of ChatGPT initially led to significant improvements in learning outcomes; however, these effects gradually diminished over time, possibly due to a decrease in novelty or the requirement for ongoing interaction and adaptation of technology to maintain engagement. Kavadella et al. (2024) employed a mixed-methods approach in a 16-week experiment, demonstrating that ChatGPT positively and significantly enhanced students’ performance on knowledge assessments. Long-term use of ChatGPT has resulted in continued enhancements in academic performance and learning habits, with notable progress in autonomous learning abilities. Kosar et al. (2024) conducted a 15-week controlled experiment involving 182 participants in an introductory undergraduate course on object-oriented programming. Results showed ChatGPT users had a slightly higher success rate in lab work than the control group, but the difference was not statistically significant. ChatGPT did not affect academic performance, practical work, or mid-term scores. Its effectiveness may depend on intervention duration and might need more time to show long-term learning effects.
The effectiveness of ChatGPT varies among students at different educational levels. Bachiri et al. (2023) conducted a study on fifth-grade students to examine the impact of gamification and automatic problem generation on K-12 student engagement and learning outcomes. The results indicated that gamification can boost engagement, motivation, and learning outcomes among K-12 students. Alneyadi and Wardat (2023) found that ChatGPT positively influenced the academic performance and learning perceptions of eleventh-grade students in Emirates schools, particularly in the field of electronic magnetism. Essel et al. (2024) recruited 125 university students and found that ChatGPT significantly and positively affected their critical, reflective, and creative thinking. Karaman (2024) explored whether ChatGPT improved mathematics scores among third-grade students preparing lesson plans for elementary mathematics courses. The study revealed that the experimental group outperformed the control group, but the difference was not statistically significant.
Chen and Chang (2024) examined seventh-grade students and observed notable differences between the GameGPT example group and a purely game-based group. Students who learned with the assistance of examples from GameGPT performed better than those who learned solely through the games. However, simply providing ChatGPT did not improve students’ understanding or their ability to solve problems. Abbas et al. (2024) surveyed university students and discovered that those experiencing higher academic workloads and time pressures were more likely to use GenAI. In contrast, reward-sensitive students were less inclined to use GenAI, and academic load, time pressure, and reward sensitivity indirectly affected student performance through their use of GenAI. Overuse of GenAI could lead to procrastination and memory decline, possibly reducing students’ academic performance. Sun and Zhou (2024) conducted a meta-analysis to examine the impact of the GenAI-supported programming mode on university students’ academic achievement and found that GenAI significantly boosted the academic performance, with a medium effect size. When the generated content was textual and the sample size was between 21 and 40, using an independent learning style resulted in the greatest improvement in students’ academic achievement. These findings indicate that educational level, learning content, and individual differences influence the effectiveness of ChatGPT, highlighting the need for further research on best practices and implementation strategies.
The review revealed many empirical reports on students using ChatGPT across educational levels, but results varied. Most research is centred on higher education, rather than secondary and elementary education, with no studies on preschool students. No systematic comparison of ChatGPT’s effects across levels exists. Future reviews should investigate whether ChatGPT impacts students differently at preschool education and other levels.
The current research demonstrates significant variations in the effectiveness of ChatGPT across different knowledge types, a critical finding that has been largely overlooked in existing meta-analyses (e.g., Zhu et al., 2025; Deng et al., 2025). Through a systematic analysis of empirical studies, we identified distinct patterns: For procedural knowledge (e.g., nursing skills, programming), Chang et al. (2024) integrated ChatGPT into nursing curriculum design and found significant improvements in learners’ critical thinking, problem-solving abilities, and learning engagement. Urban et al. (2024) further confirmed that ChatGPT can substantially enhance students’ self-efficacy and task completion quality in complex creative tasks while reducing cognitive load. However, Zhai et al. (2023) observed relatively weaker effects on knowledge transfer, possibly due to the hands-on practice needed for procedural knowledge. Regarding declarative knowledge, the research outcomes are contradictory. Alneyadi and Wardat (2024) demonstrated through quasi-experimental studies that ChatGPT significantly improves test scores, comprehension, and reasoning abilities. Zhai et al. (2023) ‘ChatGPT+metaverse’ collaborative learning model is particularly effective for rote memorisation. Conversely, Shoufan (2023) reported an overall effect size of −0.16 for computer engineering students’ declarative knowledge acquisition, suggesting potential negative impacts.
Existing research shows significant variations in ChatGPT’s effectiveness across different instructional approaches, a dimension poorly explored in current meta-analyses (e.g., Zhu et al., 2025; Deng et al., 2025). Our review found that in blended learning, Lee et al. (2024) reported that AI chatbots boosted engagement and critical thinking. Zheng et al. (2024) confirmed that ChatGPT dialogue systems enhance online collaborative learning, fostering knowledge activation and contribution. The flipped classroom showed mixed results, with Yilmaz and Yilmaz (2023) noting improvements in computational thinking, programming self-efficacy, and motivation when using generative AI tools in education.
However, Zhu et al. (2023) found no significant differences in interdisciplinary problem-solving among STEM students with or without ChatGPT integration. The other instructional modes demonstrated unique effectiveness patterns. Luo et al. (2024) discovered that generative teacher feedback improved self-regulated learning in traditional mathematics classrooms but had a limited impact on academic performance. Li et al. (2024) intelligent Socratic approach showed notable benefits for problem-solving and creative thinking, yet minimal effects on learning performance and critical thinking.
In summary, ChatGPT has significant potential to improve students’ learning outcomes, but its effectiveness varies due to factors like subject background, duration, educational level, types of knowledge, and instruction modes. This suggests further research is needed to understand these influences. Current meta-analyses do not systematically examine knowledge types and instructional modes as moderating variables (see Supplementary Table S1). This leads to three key limitations: (1) unresolved sources of heterogeneity (e.g., Laun and Wolff, 2025), (2) inaccurate effect size estimations (e.g., Deng et al., 2025), and (3) a lack of exploration into ChatGPT’s applicable educational environments (e.g., Deng and Yu, 2023; Wu and Yu, 2024). Additionally, some studies report insufficient effect sizes and limited literature inclusion (e.g., Alemdag, 2025). To address these limitations, we introduce knowledge types and instructional modes as moderator variables to clarify heterogeneity sources, provide an accurate effect size and identify optimal educational settings for ChatGPT use.
As a result, crucial questions remain unanswered: ‘Does ChatGPT significantly influence student learning outcomes?’ and ‘Do various moderating variables (such as subject, experimental duration, educational level, type of knowledge, and instructional mode) affect these outcomes, and if so, what are the differences?’ Therefore, conducting a systematic meta-analysis is essential to synthesise and integrate existing research findings and offer a more comprehensive understanding of the impact of ChatGPT on student learning outcomes. By aggregating extensive research data, we can identify patterns, trends, and effect sizes that will inform future research and shape educational practices, leading to more accurate and dependable conclusions. This approach can further validate or refine current theoretical perspectives and reveal new directions and methods for future research.
This study aimed to examine ChatGPT’s impact on student learning, providing insights for future research and educational applications. We conducted a systematic review and meta-analysis to synthesise data, assess study quality, quantify effect size, explore heterogeneity, and check for publication bias (Wu et al., 2024a). This approach sought to clarify ChatGPT’s overall effect and identify factors affecting its effectiveness.
Recent reviews and meta-analyses have explored ChatGPT in education, mainly focusing on specific subjects or education levels. While early results show positive effects of GenAI, effect sizes vary. Additionally, past studies have not thoroughly examined moderator variables influencing ChatGPT’s impact on student learning outcomes. Although prior research has contributed to the understanding of ChatGPT, there is still a lack of meta-analyses examining the effects of ChatGPT on student learning outcomes in a holistic context. This research aimed to examine ChatGPT’s impact on student learning outcomes, guiding educators, researchers, and developers in AI-assisted education. We conducted a meta-analysis to understand ChatGPT’s effects and potential moderator variables. Two research questions (RQs) were proposed.
RQ1: What is the overall effect size of ChatGPT on students’ learning outcomes?
RQ2: Which are the moderating variables—such as subject, experimental duration, educational level, type of knowledge, and instructional mode—that affect the impact of ChatGPT on student learning outcomes, and how do these variables influence its effectiveness?
This meta-analysis was designed to thoroughly examine the effects of ChatGPT on student learning outcomes. The research process followed rigorous scientific methodologies, including a systematic literature search, setting inclusion criteria, careful literature screening, detailed literature coding, rigorous validation of coding accuracy, precise calculation and summarisation of effect sizes, thorough testing for publication bias, and comprehensive heterogeneity analysis (Wu et al., 2024b). To maintain the integrity of the meta-analysis, this study adhered to the criteria established by Shelby and Vaske (2008), ensuring the reliability and accuracy of the findings.
This study employed a multi-database comprehensive retrieval strategy utilising authoritative literature databases, including the Web of Science, Wiley Online Library, SpringerLink, ProQuest Education, Elsevier Science Direct, and CNKI. The search keywords comprised ‘Generative Artificial Intelligence’ ‘GenAI’ ‘ChatGPT’ or ‘AI-generate’ in conjunction with ‘learning outcome’, ‘learning performance,’ ‘education,’ ‘course’ and ‘student’ to deeply examine the influence of ChatGPT on learning outcomes and performance within the educational domain. To ensure the quality and reliability of the literature, this study prioritised papers published in core journals. The literature search was limited to articles published between November 2022 and June 2024. After deduplication, 2038 articles were screened to provide robust foundational data for this study.
We used a thorough screening process for the retrieved literature to ensure the quality and dependability of the meta-analysis results. The inclusion and exclusion criteria were as follows: (1) the research must use ChatGPT either as a direct learning tool or as a supported learning application, with a specific focus on its impact on student learning outcomes; (2) the research methodology was limited to experimental designs, including both experimental and quasi-experimental approaches; (3) the experimental design must include the establishment of experimental and control groups or feature pre- and post-test assessments; (4) the participants must be limited to primary and secondary school students and college students, we excluded studies focused solely on special education students, adult learners, and ChatGPT itself. Our aim was to isolate the effects of ChatGPT within mainstream student populations, thereby ensuring clear interpretability of the results; and (5) the study must provide sufficient data to calculate the effect size (see Table 1). We calculated the pooled effect sizes using the post-test scores from the experimental and control groups. For articles that contained multiple studies or experiments, an effect size was calculated for each independent comparison only if it provided quantifiable outcome data. If a study solely reported qualitative findings, no effect size was extracted from it. English represents the main language of international academic exchanges in this field. To provide important insights into the localisation practice of AI educational applications in ChatGPT-restricted contexts, we supplemented the inclusion of Chinese literature. This combined approach guarantees that the GenAI meta-analysis encompasses both global trends and region-specific practices (Deng et al., 2025; Zhu et al., 2025). Strict adherence to these stringent screening criteria ensures the scientific validity and relevance of the selected literature samples and provides a strong methodological basis for this meta-analysis.
All the searched articles were managed with the Endnote X9 software. Two researchers independently screened the records and extracted data according to the inclusion and exclusion criteria. Any disagreements during this process were resolved until a consensus was reached. We initially conducted a literature search using predefined keywords within a preselected database that met the above criteria. Subsequently, we collected a total of 2038 articles after removing duplicates. After examining titles and abstracts, articles unrelated to the predefined experimental subjects were excluded, resulting in a final count of 463. Additionally, a comprehensive review of these 463 articles was performed, focusing on the standardisation of experimental designs and the completeness of the research data. After screening, 35 articles meeting the specified criteria were included. Of these, 28 were publicly accessible English-language sources (80%), and seven were Chinese-language sources (20%), yielding 134 valid research effects (see Fig. 1). To increase the transparency of the meta-analyses and ultimately improve the quality and reproducibility of the research, we selected studies using the PRISMA guidelines and selection procedures (Page and Moher, 2017).
PRISMA study flow phases.
The quality of the included studies was assessed using a structured checklist adapted from Montuori et al. (2023), which has been applied in a prior meta-analysis. Given the diversity of sources, a consistent quality check is essential to mitigate the risk of including low-quality studies. The checklist evaluates seven key aspects: quality of the experimental design, presence of an active control group, use of pre–post-test measures, reliability of the assessment tool, pre-test equivalence, clarity of the method, and clarity of the intervention (see Supplementary Table S2). Each quality indicator was scored as 1 (met) or 0 (not met), with a maximum total score of 7, and higher scores indicating better research quality. The studies included in our meta-analysis achieved an average score of 6.11 (SD = 1.12), and the score for the Chinese studies was 6.4 (SD = 0.79). There was no statistically significant difference in quality scores between the Chinese and English studies (t = 0.82, df = 34, p = 0.42), suggesting that they were of relatively high quality and suitable for inclusion in our analysis. This quality check was vital to ensure the reliability of our findings, particularly considering the potential variability in source quality.
The selected studies underwent detailed coding to facilitate the statistical analysis and calculation of effect sizes. The coding encompassed basic information about the studies (e.g., article title, author names, and publication year) and substantive content (e.g., sample size, overall effect size, and moderating variables). The moderating variables included subjects, experimental duration, educational level, instructional mode, and type of knowledge, as they play a vital role in explaining the variability in effect sizes. The encoding conventions were as follows (see Supplementary Table S3): sample size coded as S (indicating a small sample size, defined as fewer than 100 participants) or L (indicating a large sample size, defined as 100 or more participants); subjects coded in experimental research as physics, chemistry, English, mathematics, computer science, educational technology, literature, science, and STEM, with subjects having smaller sample sizes categorised as ‘other’; experimental duration coded in a standardised way based on the length of experiment intervention (Laun and Wolff, 2025); educational level coded as primary educatio (PE), secondary education (SE), and higher education (HE) according to the student grade level; instructional mode refers to the structural integration of ChatGPT into teaching. Instructional mode coded on the basis of the type of instructional mode, with the traditional instructional mode denoted as Le; other modes including online collaborative learning (OCL), flipped classroom approach (FCA), problem-based learning (PBL), game-based learning (GBL), blended learning (BL), cooperative learning (CL), and online learning (OL), which we code as innovative instructional modes; we classify types of knowledge coded as DK (Declarative Knowledge) and PK (Procedural Knowledge) according to Anderson and Krathwohl (2001). Declarative knowledge involves understanding facts, concepts, and information (Anderson and Krathwohl, 2010; Ryle and Tanney, 2009). Procedural knowledge pertains to skills and procedures demonstrated through action, emphasising the application of learned skills in practical contexts (Ryle and Tanney, 2009). The coding process was independently carried out by two researchers skilled in meta-analytic coding, who communicated with each other to ensure a consistent and reliable understanding of each moderator. Coding consistency was assessed using the kappa coefficient. The resulting kappa value was 0.851, demonstrating high coding consistency and reliability of the screening criteria.
This study used the Comprehensive Meta-Analysis software (version 3.0) to carry out the meta-analysis. Cohen’s d can produce biased estimates, particularly for small sample sizes. Therefore, we adopted Hedges’g as the standardised effect size (ES) metric to assess effect sizes more accurately. The analysis included effect size calculation, heterogeneity testing, publication bias assessment, and examination of the moderating variables’ impact. To ensure the reliability and accuracy of the results, a comprehensive effect test was performed, encompassing both publication bias assessment and heterogeneity testing. Publication bias was evaluated using three methods: generating a funnel plot, calculating the classic fail-safe number, and applying Begg’s test to detect the presence of publication bias. Heterogeneity testing utilised Q and I2 statistics to assess study variability and identify potential biases related to publication or research. Additionally, through a moderator effect test, we examined whether various individual factors (subject, experimental duration, educational level, instructional mode, and type of knowledge) moderated the impact of ChatGPT on student learning outcomes. Following thorough testing and analysis, the research findings provide empirical support for understanding the effects of ChatGPT and highlight its main influencing factors and potential moderating effects.
The study comprehensively examined the effects of ChatGPT on student learning outcomes, supported by a detailed research framework. To address the first research question, we synthesised data from 35 educational studies on ChatGPT’s impact on student learning outcomes, encompassing 134 effect sizes. We systematically assessed the overall effect size of ChatGPT on student learning outcomes and tested for publication bias and heterogeneity. We examined the moderating roles of various individual factors, including subject, experimental duration, educational level, instructional mode, and type of knowledge in educational experiments. These moderating variables significantly influenced the effectiveness of ChatGPT on student learning outcomes and are essential elements in empirical research.
This study employed Cohen’s (1992) criteria to evaluate the effect size, with effect sizes of approximately 0.2, 0.5, and 0.8 denoting small, moderate, and large effects, respectively. The research findings indicated that the mean effect size of ChatGPT on student learning outcomes was 0.670, with a p-value of 0.000, which was statistically significant. The combined effect size fell within the range of 0.5–0.8, suggesting that the impact of ChatGPT on student learning outcomes was moderate (see Table 2). The 95% confidence intervals computed using the random-effects model were [0.495 and 0.844], excluding zero, which further substantiated that the effect of ChatGPT on student learning outcomes was not due to chance. Among the 35 studies included, only three exhibited negative effect sizes. Most studies suggest that ChatGPT positively influenced student learning outcomes, significantly enhancing learning outcomes.
Following the classification method proposed by Gu and Hu (2018), this study categorised student learning outcomes into cognitive and non-cognitive dimensions. The cognitive dimension includes creativity, social skills, problem-solving ability, learning achievement, and critical thinking, while the non-cognitive dimension covers interest in learning, engagement, motivation, and self-efficacy. Statistical analysis revealed that the effect sizes for both the cognitive (g = 0.872, p < 0.05) and noncognitive (g = 0.539, p < 0.05) dimensions exceeded 0.5, which was statistically significant, indicating that ChatGPT exerts a positive and significant influence on cognitive and non-cognitive learning outcomes (Table 2). The effect size for the cognitive dimension exceeded 0.8, indicating a significant positive impact of ChatGPT on students’ cognitive abilities. The effect size for the non-cognitive dimension was ~0.5, indicating a moderate positive and significant effect of ChatGPT on students’ non-cognitive aspects. The heterogeneity test results revealed a Q-value of 7.337 and a p-value of 0.007, which were <0.05, demonstrating a significant difference between the two types of learning outcomes. This illustrates that ChatGPT has varying degrees of effect on students’ cognitive and noncognitive learning outcomes. In summary, the enhancing effect of ChatGPT on student learning outcomes was comprehensive, effectively improving both cognitive and noncognitive learning outcomes, with a more pronounced improvement in the cognitive dimension.
Regarding cognitive dimensions, this study identified five learning outcomes: creative thinking (n = 7), social skills (n = 2), problem-solving ability (n = 10), learning achievement (n = 29), and critical thinking (n = 8). Table 2 provides detailed information on the learning outcomes for the five cognitive dimensions. The effect size of creative thinking (g = 0.633, p < 0.05) was above 0.5 and statistically significant, indicating that ChatGPT effectively enhanced students’ creative thinking. The effect size of the social skills dimension was 0.481 (p < 0.05), indicating that ChatGPT had a moderately positive and significant effect on improving students’ social skills. The effect sizes of problem-solving ability (g = 0.933, p < 0.05), learning achievement (g = 0.876, p < 0.05), and critical thinking (g = 1.008, p < 0.05) were all above 0.8 and statistically significant, suggesting that ChatGPT had a significant positive effect on these dimensions, especially on the development of critical thinking. The heterogeneity test results (Q = 5.454, p = 0.244 > 0.05) indicated no significant differences across cognitive dimensions.
Regarding non-cognitive dimensions, this study primarily identified four learning outcomes: learning interest (n = 3), self-efficacy (n = 9), learning engagement (n = 7), and learning motivation (n = 10). Table 2 shows the learning outcomes for the four non-cognitive dimensions. The effect size of learning interest (g = 0.915) exceeded 0.8, indicating a high effect size. However, the p-value of 0.085 was greater than the critical value of 0.05. This may be related to the insufficient research design and sample size, resulting in a lack of statistical significance. The effect sizes of self-efficacy (g = 0.467, p < 0.05), learning engagement (g = 0.660, p < 0.05), and learning motivation (g = 0.409, p < 0.05) were all approximately 0.5 and statistically significant, indicating that ChatGPT had a moderate positive effect on these dimensions, with the most pronounced effect on fostering students’ learning engagement. The heterogeneity test results (Q = 2.820, p = 0.420 > 0.05) showed no significant differences in effects across the non-cognitive dimensions, indicating that ChatGPT positively enhanced all four non-cognitive dimensions.
To evaluate whether small-sample studies disproportionately influenced our results, we conducted leave-one-out sensitivity analyses, paying special attention to studies with k < 3 (see Supplementary Table S4). The pooled effect size remained stable (range: 0.628–0.696 vs. original 0.670, 95% CI [0.495, 0.844]). The exclusion of two small-sample studies (N < 150) caused a maximum deviation of 6.3%, and no exclusion altered the statistical conclusions. This confirms that our findings are not driven by limited sample data. In terms of the subgroup analysis, although the primary school group contained two effect sizes that resulted in a wide confidence interval (ES = 0.824, 95% CI [−0.198,1.846]) and the results were not significant, its weight proportion was only 5%, which was much lower than that of the dominant middle/university group (weight 95%). After further removing primary school group A from the entire model, its effect size decreased to 0.280 (95% CI [−0.369, 0.929]), and its weight increased to 8%. The intergroup heterogeneity test was still not significant (Qbet = 1.217, p = 0.270), indicating that the presence or absence of a primary school group had no substantial impact on the judgement of subgroup differences. However, excluding Study B significantly increased the effect size to 1.329 (95% CI [0.896,1.762]), with a weight proportion of 15% and statistically significant differences between the groups (Qbet = 7.884, p = 0.005) (see Supplementary Table S5). These results fully demonstrate that, although there is a problem of insufficient statistical power in the primary school subgroup due to the limited sample size, the main conclusions of this study are robustly driven by the large sample size of the middle and university groups and have high methodological robustness.
Publication bias is a common error in meta-analyses that can distort results. This study used three methods to evaluate and reduce the bias. First, a funnel plot was used for testing. The points representing the effect sizes of the study sample were evenly distributed on both sides of the central axis, where the mean effect sizes were located, with most effect sizes concentrated in the upper-middle part of the funnel plot. This suggests a low probability of the meta-analytical findings being influenced by publication bias (Fig. 2). Second, to avoid potential bias from visual inspection, this study applied the classic fail-safe number for evaluation with a measurement standard of 5n + 10, where n denotes the number of studies included in the meta-analysis. The classic fail-safe number for this study was 3522, which was significantly higher than the critical value, indicating that the estimated effect size of the unpublished research was unlikely to have a substantial impact on the effect size of the meta-analysis. Finally, Begg’s test for funnel plot asymmetry and the classic fail-safe N test were used as complementary procedures to investigate potential publication bias. The results indicated a z-value of 1.757, which was <1.96, and a p-value of 0.079, which exceeded 0.05, further confirming the absence of publication bias. In summary, these three methods showed no significant publication bias, and the effect sizes were stable. The conclusions accurately reflect the actual situation.
Funnel plot.
According to the statistical principles of meta-analysis, data homogeneity is a prerequisite. Therefore, it is essential to perform heterogeneity tests based on the results of multiple studies. Commonly used heterogeneity testing methods include the Q-test and I2 tests. The Q-test assumed homogeneity across all studies and primarily focused on p-values. If the p-value exceeds 0.1, it indicates homogeneity in the study; otherwise, it indicates heterogeneity. The I2 statistic reflects the proportion of heterogeneity in the total variation of the effect sizes, with values ranging from 0 to 100. A larger I2 value indicates greater heterogeneity between samples. Higgins et al. (2003) suggest using 25%, 50%, and 75% as critical values for low, medium, and high heterogeneity, respectively. The heterogeneity test results of this study, as shown in Table 3, yielded a Q value of 409.067 with a p-value of 0.000, which was <0.01, indicating significant heterogeneity among the samples. The I2 value was 91.444%, indicating that 91% of the variation was due to true differences in effect sizes. Owing to the heterogeneity of the research sample, this study utilised a random-effects model to consider intra- and inter-study variability to conduct meta-analysis effect size tests (Wilson and Lipsey, 2001). Therefore, this study employed a random-effects model to evaluate the effects of ChatGPT on student learning outcomes.
This study analysed 35 articles covering subjects including physics (n = 2), chemistry (n = 1), English (n = 8), mathematics (n = 3), computer science (n = 4), teaching skills (n = 3), literature (n = 3), interdisciplinary (n = 1), science (n = 1), and more than 10 other subjects. The meta-analysis results indicate that ChatGPT has been studied and applied across many subjects; however, current research on ChatGPT applications predominantly focuses on disciplinary education. The sample size was relatively small. Statistical results revealed that ChatGPT had a significant positive effect (ES > 0.8) on subjects such as physics (g = 1.951, p < 0.05), chemistry (g = 1.276, p < 0.05), and English (g = 0.994, p < 0.05), suggesting that ChatGPT had a strong promoting effect on these subjects (see Table 4). ChatGPT exhibited varying degrees of moderate positive effects (ES = ± 0.5) on subjects such as mathematics (g = 0.655, p < 0.05), computer science (g = 0.436, p < 0.05), teaching skills (g = 0.421, p < 0.05), and literature (g = 0.376, p < 0.05), indicating that ChatGPT had a positive effect on learning across subjects, but its effectiveness can be improved. The experimental results for interdisciplinary (g = 0.16, p > 0.05) and science (g = 0.169, p > 0.05) subjects did not reach a statistically significant level and thus lacked statistical significance. This may be due to an insufficient research design and sample size. Further research is needed with respect to ChatGPT. According to the heterogeneity test results (Q = 93.933, p = 0.000 < 0.05), there were significant differences in the effect of ChatGPT on student learning outcomes across subjects, implying that its impact on learning outcomes varied unevenly across subjects.
According to the coding table, this study categorised the experimental duration into three lengths:‘less than one month’ (n = 7), ‘one month to three months’ (n = 18), and ‘more than three months’ (n = 7). The statistical results indicated a moderate positive effect (ES = ±0.5) of ChatGPT on the experimental durations of ‘less than one month’ (g = 0.456, p < 0.05) and ‘one month to three months’ (g = 0.669, p < 0.05); for the experimental intervention of ‘more than three months’(g = 1.033, p < 0.05), there was a significant positive effect (ES > 0.8; see Table 4). According to the heterogeneity test results (Q = 9.204, p = 0.01 < 0.05), there were significant differences in the effect of ChatGPT on student learning outcomes across experimental durations, indicating that this effect varied unevenly across experimental durations, with students experiencing an experimental duration of more than three months exhibiting the best learning outcomes.
This study categorises educational levels into primary, secondary, and higher education based on the sample data. This study predominantly focused on higher education (n = 26), followed by secondary education (n = 7); primary education was the least represented (n = 2). The statistical results indicate that ChatGPT’s effect size on primary education students’ learning outcomes was 0.824, which falls within the high effect size category. However, as the p-value of 0.114 was significantly larger than the critical value of 0.05, the difference was not statistically significant. To ensure accurate interpretation, we excluded the primary education group. Subsequent sensitivity analysis revealed that all excluded effect sizes were comparable to the original merged values, and their overlapping confidence intervals further confirmed the robustness of the results against individual study biases. ChatGPT exhibited a significant positive effect (ES > 0.8) on the learning outcomes of students in secondary education (g = 0.847, p < 0.05) and a moderate positive effect (0.5 < ES < 0.8) on the learning outcomes of students in higher education (g = 0.744, p < 0.05; see Table 4). According to the heterogeneity test results (Q = 0.399, p = 0.527 > 0.05), there was no significant difference in the effect of ChatGPT on the learning outcomes of students across educational levels. This indicates that ChatGPT had a positive effect on improving the learning outcomes of students at various educational levels, with the most significant effect observed among secondary education students.
Guided by certain teaching theories, this study categorised instructional modes into traditional (n = 19) and innovative (n = 16). The traditional instructional mode is exemplified by the traditional lecture method, whereas the innovative instructional mode encompasses various instructional methods, such as interdisciplinary and gamified teaching and flipped classroom approaches. The statistical results indicated that the effect size of ChatGPT on the traditional instructional mode (g = 0.928, p < 0.05) exceeded 0.8, thereby demonstrating a significant positive effect of ChatGPT on the traditional instructional mode. The effect size of the innovative instructional mode (g = 0.623, p < 0.05) ranged between 0.5 and 0.8, also achieving a significant level and implying that ChatGPT has a moderate positive effect on the innovative instructional mode (see Table 4). According to the heterogeneity test results (Q = 4.080, p = 0.043 < 0.05), ChatGPT exhibited significant differences across instructional modes, indicating that it positively enhanced student learning outcomes in various instructional modes. Students’ learning outcomes in the traditional instructional mode were superior to those in the innovative mode.
This study examined the effects of ChatGPT on student learning outcomes for various knowledge types. The statistical results showed that the effect size of ChatGPT on student learning outcomes for declarative knowledge (n = 13) was 0.903, reaching a significant level (p < 0.05), thereby demonstrating the notable positive impact of ChatGPT on declarative knowledge. The effect size of procedural knowledge on student learning outcomes (n = 22) was 0.705, which was also significant (p < 0.05), indicating that ChatGPT had a moderately positive effect on procedural knowledge (see Table 4). According to the heterogeneity test results (Q = 1.599, p = 0.206 > 0.05), there was no significant difference in the efficacy of ChatGPT across different types of knowledge, suggesting that ChatGPT had a positive effect on students’ learning of declarative and procedural knowledge; the student learning outcomes of declarative knowledge were particularly superior to those of procedural knowledge.
The results showed that the ChatGPT application significantly improved student learning outcomes, demonstrating a moderate positive effect (g = 0.670, p < 0.001). This finding highlights the potential of ChatGPT to enhance educational quality and boost students’ learning efficiency. Subsequently, we investigated various moderating variables that affected the effectiveness of ChatGPT, including academic subjects, experiment duration, educational levels, instructional modes, and types of knowledge. The analysis revealed that ChatGPT effectively improves student learning outcomes across different subjects, educational stages, and instructional modes. This finding offers valuable guidance for educators in refining the application strategies of ChatGPT to suit specific contexts, thereby maximising educational benefits. The meta-analysis results reaffirm the importance of ChatGPT in education and reveal its broad applicability and positive effects in diverse educational settings and conditions. These findings carry significant theoretical and practical implications for advancing innovation and practice in education technology.
We have supplemented our discussion with recently published literature. Arkan et al. (2025) validated that training provided on the use of ChatGPT in the nursing process positively affected nursing students’ problem-solving skills, attitudes towards artificial intelligence, competencies related to the nursing process, and levels of satisfaction. The results of Urban et al. (2025) indicated that ChatGPT assistance enhances the integration of correct human- and AI-generated content. Zhou et al. (2025) demonstrated that the ChatGPT-facilitated scaffolding significantly improves students’ mathematical problem-solving performance. There was no significant difference between the recently published high-quality literature results recently published and those of our study.
Based on a meta-analysis of 35 relevant published articles encompassing 134 effect sizes, this study revealed that ChatGPT had a significantly and moderately positive effect on student learning outcomes (g = 0.670, p < 0.001). This finding underscores the ChatGPTs positive role of ChatGPTs in enhancing learning outcomes in various educational settings. Analysis shows ChatGPT boosts both cognitive and noncognitive skills in students, with a stronger effect on cognition. It improves creative thinking, problem-solving, learning achievement, and critical thinking, supporting cognitive growth. It also enhances engagement, self-efficacy, and motivation, fostering noncognitive skills. Bias tests found no significant bias, confirming the robustness of these findings despite ChatGPT’s recent appearance in education technology.
The comprehensive application of ChatGPT across all educational levels, subjects, and processes holds promise for improving student abilities across multiple dimensions. ChatGPT-assisted teaching has the potential to create personalised, adaptive, and engaging learning environments tailored to individual student needs, thereby revolutionising educational practices (Chen and Chang, 2024). Future research should explore the long-term effects of ChatGPT in education and investigate how its integration can be optimised across educational contexts. Further studies are required to address the potential challenges and ethical considerations associated with the widespread adoption of ChatGPT in educational settings.
This study shows that ChatGPT positively impacts student learning, improving cognitive and non-cognitive skills. Findings indicate ChatGPT is a promising tool for educational change and merits further exploration and use practices.
To better understand the conditions under which ChatGPT is effective, we examined five moderating variables relevant to how ChatGPT enhances student learning outcomes: subject, experimental duration, educational level, instructional mode, and type of knowledge. Specifically, while educational level and type of knowledge did not show a statistically significant effect on learning outcomes (p > 0.05), subject, experimental duration, and instructional mode were found to have a positive and significant influence (p < 0.05). These findings indicate that integrating ChatGPT into teaching practices may be particularly effective when tailored to specific subjects, maintained over a longer duration, and implemented within an interactive instructional mode.
At the educational level, the analysis of between-group effects revealed that ChatGPT did not exhibit a significant differential effect on student learning outcomes across the various stages (Qbet = 0.399, p = 0.527 > 0.05). The results indicate that ChatGPT positively influences learning outcomes at the primary, secondary, and tertiary levels (Xing et al., 2025; Kohnke et al., 2025), which is inconsistent with the conclusions of Wang and Fan (2025). Notably, the relationship between these outcomes and educational level was nonlinear; there was no monotonic increase or decrease with the progression of educational levels. In contrast, ChatGPT technology in secondary education demonstrated a notably more effective and significant effect, suggesting that this stage is pivotal for mastering and applying ChatGPT in innovative learning activities (Alemdag, 2025; Wu and Li, 2024; Liu et al., 2025). Although the sensitivity analysis supports the robustness of the main conclusions, the results of small-sample subgroups (e.g., primary education group) need to be interpreted with caution. Possible sources of bias in primary education group B may arise from sample characteristics, differences in intervention methods, and other factors. GenAI can provide answers and examples for conceptual understanding and knowledge points, making it easier for students to grasp (Giannakos et al., 2024). This learning content are relate to the accumulation of knowledge and academic performance in K-12 education. In contrast, higher education content is more complex and diverse, making it difficult to fully understand by merely mastering individual concepts, let alone fostering students’ practical skills and promoting the development of their higher-order thinking (Biggs et al., 2022; Dasari et al., 2024). Additionally, researchers have found that younger students are more familiar with robots (Fernández-Llamas et al., 2018). Secondary school students exhibit greater curiosity and anticipation of AI technologies such as ChatGPT, coupled with increased enthusiasm and exploratory motivation for learning. When K-12 learners engage in simpler tasks, they are more attracted by the entertaining features of AI chatbots, which may enhance ChatGPT’s impact on their learning outcomes (Wu and Li, 2024). Therefore, exploring teaching models, educational frameworks, and teaching agents suited to different levels of education is highly significant for the high-quality development of future education.
At the knowledge level, ChatGPT demonstrated a consistently positive effect on improving student learning outcomes across different types of knowledge, with no significant differences observed among them (Qbet = 1.599, p = 0.206 > 0.05). This consistency indicates that, despite the distinct characteristics of various knowledge types, ChatGPT’s influence remains relatively uniform. Although no significant differences were identified, the findings suggest that ChatGPT slightly enhances students’ declarative knowledge learning outcomes more than procedural knowledge outcomes (Yuxian, 2025). ChatGPT supports declarative knowledge instruction by generating texts, images, sounds, and videos, and by using real-life scenarios or personal experiences as learning contexts (Wang et al., 2024). It excels at breaking down complex ideas into well-organised and accessible knowledge, thereby reducing extraneous cognitive load and deepening learners’ understanding (Dasari et al., 2024), making it particularly suitable for conveying declarative knowledge. Procedural knowledge, however, relies heavily on practical environments that require hands-on practice and dynamic adjustments, which are challenging for ChatGPT to effectively simulate.
Without a structured framework in the ChatGPT learning environment, students may struggle to integrate procedural knowledge from different fields, leading to an inconsistent understanding and application of knowledge (Wang et al., 2023). This aligns with previous findings that AI-driven teaching tools typically require scaffolding to support higher-order cognitive processes (Chi and Wylie, 2014). In the AI era, the classification of knowledge should focus on the intrinsic connections among its various types. Declarative knowledge forms the foundation for the application of procedural knowledge, and all types of knowledge function as interconnected wholes (Lohse and Healy, 2012; Chen et al., 2025). A large language model can improve the learning of declarative knowledge, but its effect on procedural learning is limited. Critical thinking guidance can support procedural learning, but it increases cognitive load. Combining both approaches creates a synergistic effect, with critical thinking activating procedural learning and large language model chatbots reducing cognitive load, thus distributing cognitive resources more evenly (Yuxian, 2025). Although ChatGPT is seen as enhancing learning efficiency and productivity, it should not be used indiscriminately but carefully tailored to the specific nature of different types of knowledge (Mogavi et al., 2023). Educators must consider different types of knowledge in instructional design. Ignoring this may reduce its effectiveness. Additionally, learners should critically assess the accuracy of AI-generated content (Rees, 2022).
At the subject level, there were significant differences in the effects of ChatGPT on student learning outcomes across different subjects (Qbet = 93.933, p = 0.000 < 0.05). ChatGPT has a notably positive impact on most subjects, especially physics, chemistry, and English, whereas its effect on interdisciplinary and scientific subjects is less evident. The influence of ChatGPT on students’ learning outcomes showed statistically significant variations across subjects, contradicting the findings of Deng and Yu (2023) and Alemdag (2025). This can be attributed to the chosen classification criteria. However, some researchers support this perspective (Zhu et al. 2025). This variation could be due to several factors. Jeong et al. (2019) found that students’ disciplinary backgrounds influence how technology affects learning outcomes. Students in natural sciences and language tend to see ChatGPT as a helpful tool because they often engage with technology and rely on tools for coding and data analysis (Jin et al. 2025). Conversely, Humanities students are less likely to see ChatGPT as technical support and more likely to view it as a potential threat, fearing that it may oversimplify or misunderstand the complex aspects of human culture, ethics, and behaviour (Luo, 2025). In STEM and related courses, students are often required to collaborate with ChatGPT to complete complex projects (Li, 2023; Huesca et al., 2024), single-subject curricula are more focused, and using intelligent tools and devices in subjects with more procedural content yields more optimal cognitive outcomes for students, resulting in a more pronounced enhancement in their learning achievements (Wu, 2023). Therefore, future research should develop more refined settings for ChatGPTs based on subject characteristics. ChatGPT use needs to be optimised by considering the topics being taught and the disciplinary backgrounds of students, rather than applying it uniformly (Zhu et al., 2023).
At the experimental duration level, ChatGPT positively influenced student learning outcomes across various experimental cycles, with significant differences between them (Qbet = 9.204, p = 0.01 < 0.05). The results suggest a linear relationship between the experimental duration and its effect on student learning outcomes. This research revealed that the longer the experimental duration, the more robust the effect on student learning outcomes (Laun and Wolff, 2025; Liu et al., 2025; Zhu et al., 2025). However, researchers have different views on this (Wu and Yu, 2024; Deng et al., 2025; Wang and Fan, 2025). Some studies have suggested that short-term interventions are beneficial for improving educational technology learning outcomes (Villena-Taranilla et al., 2022) and that long-term use may lead to the disappearance of these effects (Jeno et al., 2019). The novel effects of educational technology can increase students’ interest, motivation, and engagement. However, once they use these technologies for a long time, they may disappear (Haristiani, 2019). Our research found that students who used ChatGPT for three months or longer achieved the best learning outcomes. This phenomenon may occur because interactions between learners and ChatGPT are often passive or superficial in the short term, leading to suboptimal improvements in learning outcomes. ChatGPT is not yet ready to be relied upon by students who do not have a sufficient background to evaluate the generated answers (Shoufan, 2023). Inexperienced students do not know how to initiate interactions with ChatGPT, and their prompts are too vague to guide conversations (Urban et al., 2024). This indicates that improving students’ educational technology learning outcomes cannot rely solely on interest and participation. Therefore, we must pay attention to students’ prior knowledge, self-regulation abilities, and AI literacy (Urban et al., 2024; Ng et al., 2024; Seo et al., 2021). However, as the duration increases, learners have more time to adapt to technology, improve their proficiency, and subsequently improve their performance (Liu et al., 2025). Students require a period of adjustment to integrate new tools effectively into their learning practices. The results align with Venkatesh et al. (2003)’s ‘technology adoption process,’ showing extended trial and ongoing enhancement of students’ tech skills. Increased proficiency led to better learning outcomes. Educators and researchers should focus on long-term implementation to maximise ChatGPT benefits (Lo et al., 2024; Wang and Fan, 2025). Although our study analysis confirms the moderating effect of experimental duration, two methodological limitations must be pointed out: firstly, there are significant differences in the definition of ‘long-term’ among studies (ranging from 3 to 12 months); secondly, cultural factors may affect the rate of technological adaptation. These findings suggest that future research should strive towards establishing a standardised measurement framework for technology exposure duration and conduct comparative experiments to control for cultural variables.
At the instructional mode level, ChatGPT positively improved student learning outcomes across different instructional modes, with notable differences between them (Qbet = 4.080, p = 0.043 < 0.05). This finding addresses a critical gap in prior meta-analyses (Wu and Yu, 2024; Liu et al., 2025), which reported high heterogeneity (I2 > 75%) but did not examine potential moderators. ChatGPT must be embedded, integrated, and combined with an instructional method (Weidlich et al., 2025). Without considering instructional modes, results may not apply to other educational settings. Traditional instructional modes (teacher-centred and structured) demonstrate superior compatibility with ChatGPT, as its knowledge-supplementing function seamlessly aligns with systematic pedagogy (Acerbi and Stubbersfield, 2023). Innovative instructional approaches (student-centred, exploratory) tend to yield diminished returns, probably because ChatGPT’s supportive (rather than leading) role cannot replace teacher guidance in complex tasks (Hauk and Gröschner, 2022). The strong socialisation process that students experience by sitting in traditionally oriented classrooms year after year means they are more likely to feel comfortable in environments aligned with this traditional approach, rather than in those that are cognitively more challenging (Furtak and Kunter, 2012). The experimental results of Darvishi et al. (2024) suggest that supplementing AI assistance with self-regulated strategies does not significantly outperform reliance solely on AI assistance. Similarly, Hauk and Gröschner (2022) demonstrated that unstructured, student-led activities are less effective than teacher-guided inquiries.
Students in low cognitive autonomy-supportive conditions learn significantly more, perceive significantly more choices, and rate instruction as more positive than students in high cognitive autonomy supportive conditions (Furtak and Kunter, 2012; Lazonder and Harmsen, 2016). The teacher’s central role is especially important in feedback mechanisms. Students receiving instructor feedback show markedly greater improvements in their scores than those receiving AI feedback, even after accounting for initial knowledge levels (Tian and Zhou, 2020). This advantage arises from instructors’ deep understanding of the course content, learning objectives, and assessment criteria, which allows them to provide contextually relevant and accurate feedback (Er et al., 2025). AI, despite its power, lacks contextual awareness and may struggle to grasp the nuances of assignments or rubrics, leading to less precise feedback (Liu et al., 2023). These outcomes highlight the value of traditional instructional modes. Guided by teachers, this approach provides emotional support, fosters positive emotions, and enables targeted interventions during human-computer collaboration (Baidoo-Anu and Ansah, 2023). These findings shift the focus from whether ChatGPT works to when it works best. Its compatibility with traditional methods can reduce barriers, promote integration, and support educational equity, benefiting a wide range of learners (Khowaja et al., 2024). With regard to practical implications, the grand mean effect size reported in this meta-analysis suggests that policymakers and educators should maintain or strongly consider teacher-centred traditional lecturing when using ChatGPT for instruction.
This meta-analysis shows that ChatGPT significantly improves student learning outcomes (g = 0.670, p < 0.001), boosting cognitive skills and non-cognitive skills. It has strong effects in physics, chemistry, and English, but less so in scientific and interdisciplinary fields. Optimal results occur with interventions longer than three months, especially in middle school. Combining ChatGPT with traditional methods greatly enhances learning. Its impact is greater on declarative than procedural knowledge. While effective, ChatGPT’s success depends on the subject area, duration, and instructional modes.
Although this study employed a large-scale meta-analytic approach to systematically evaluate the effect of ChatGPT on student learning outcomes, it has some limitations. First, the data mainly originated from published, peer-reviewed literature, which may have introduced publication bias. Excluding unpublished studies, which are often difficult to include in analysis, could skew results towards positive findings, reducing the comprehensiveness and objectivity of the conclusions. Second, this study may not have fully examined all the multidimensional factors influencing student learning outcomes in ChatGPT applications. While variables such as subject area, experiment duration, instructional mode, and knowledge type were considered, other unidentified or insufficiently considered factors might also significantly influence ChatGPT’s effectiveness, including but not limited to intervention setting, ChatGPT’s role and the teacher’s role. A third limitation is that the small sample size (k = 2) of the primary school group may lead to limited statistical power, and further research on the elementary education stage is necessary (Villena-Taranilla et al., 2022). Fourth, our search timeframe was limited, which could have potential impacts; future research should expand the observation period. Finally, this study mainly focused on ChatGPT’s effects on basic learning outcomes, such as improving cognitive and non-cognitive skills. It paid less attention to its impact on students’ AI literacy, computational thinking, ethical issues (e.g., data privacy and algorithmic bias), and other potential learning domains. This limitation hampers a comprehensive understanding of ChatGPT’s capacity to promote holistic student development. To strengthen methodological rigour, future studies could incorporate advanced techniques like meta-regression or three-level meta-analysis. This would provide a more nuanced interpretation of results while increasing the study’s innovative value.
To address the limitations of this study, future studies should adopt measures to enhance comprehensiveness and depth. First, researchers should broaden data sources to include peer-reviewed literature, unpublished studies, conference papers, technical reports, and relevant material. This approach reduces publication bias and enhances research objectivity and comprehensiveness of findings. Second, a multivariate analysis offers a clearer understanding of ChatGPT’s effects across contexts and their mechanisms. Future research should identify factors influencing its effectiveness, such as student differences, teachers’ instructional styles, and educational environments. Finally, future research should expand on ChatGPT’s role in improving students’ AI literacy and skills. AI in education requires clear guidelines and ethics, especially around data privacy, bias, and decision-making affecting academic paths. It should evaluate cognitive and non-cognitive skills, exploring how ChatGPT promotes critical, creative, and problem-solving abilities, supporting personalised and lifelong learning abilities.
The datasets generated and analysed during the current study, including all coding sheets, extracted data, and analysis scripts, are available as Supplementary Information files with this article.
Abbas M, Jam FA, Khan TI (2024) Is it harmful or helpful? Examining the causes and consequences of generative AI usage among university students. Int J Educ Technol High Educ 21(1):10. https://doi.org/10.1186/s41239-024-00444-7
Article Google Scholar
Abramski K, Citraro S, Lombardi L, Rossetti G, Stella M (2023) Cognitive network science reveals bias in GPT-3, GPT-3.5 turbo, and GPT-4 mirroring math anxiety in high-school students. Big Data Cogn Comput 7(3):124. https://doi.org/10.3390/bdcc7030124
Article Google Scholar
Acerbi A, Stubbersfield JM (2023) Large language models show human-like content biases in transmission chain experiments. Proc Natl Acad Sci USA 120(44):e2313790120. https://doi.org/10.1073/pnas.2313790120
Article PubMed PubMed Central CAS Google Scholar
Albdrani RN, Al-Shargabi AA (2023) Investigating the effectiveness of ChatGPT for providing personalized learning experience: a case study. Int J Adv Comput Sci Appl 14(11). https://doi.org/10.14569/ijacsa.2023.01411122
Alemdag E (2025) The effect of chatbots on learning: a meta-analysis of empirical research. J Res Technol Educ 57(2):459–481. https://doi.org/10.1080/15391523.2023.2255698
Article Google Scholar
Alneyadi S, Wardat Y (2023) ChatGPT: Revolutionizing student achievement in the electronic magnetism unit for eleventh-grade students in Emirates schools. Contemp Educ Technol 15(4):ep448. https://doi.org/10.30935/cedtech/13417
Article Google Scholar
Alneyadi S, Wardat Y (2024) Integrating ChatGPT in grade 12 quantum theory education: an exploratory study at Emirate school (UAE). Intelligence 2(4). https://doi.org/10.18178/ijiet.2024.14.3.2061
Anderson LW, Krathwohl DR (2001) A taxonomy for learning, teaching, and assessing: a revision of Bloom’s taxonomy of educational objectives: complete edition. Addison Wesley Longman, Inc
Anderson, L. W., & Krathwohl, D. R. (2010). A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives. Longman
Arkan B, Dallı ÖE, Varol B (2025) The impact of ChatGPT training in the nursing process on nursing students’ problem-solving skills, attitudes towards artificial intelligence, competency, and satisfaction levels: single-blind randomized controlled study. Nurse Educ Today 106765. https://doi.org/10.1016/j.nedt.2025.106765
Arthur F, Salifu I, Abam Nortey S (2024) Predictors of higher education students’ behavioural intention and usage of ChatGPT: the moderating roles of age, gender and experience. Interact Learn Environ 1–27. https://doi.org/10.1080/10494820.2024.2362805
Babalola VT, Ahmad SS, Tafida HS (2024) ChatGPT in organic chemistry classrooms: analyzing the impacts of social environment on students’ interest, critical thinking and academic achievement. Int J Educ Teach Zone 3(1):60–72. https://doi.org/10.57092/ijetz.v3i1.155
Article Google Scholar
Bachiri YA, Mouncif H, Bouikhalene B (2023) Artificial intelligence empowers gamification: optimizing student engagement and learning outcomes in e-learning and MOOCs. Int J Eng Pedagog 13(8). https://doi.org/10.3991/ijep.v13i8.40853
Baidoo-Anu D, Ansah LO (2023) Education in the era of generative artificial intelligence (AI): understanding the potential benefits of ChatGPT in promoting teaching and learning. J AI 7(1):52–62. https://doi.org/10.61969/jai.1337500
Article Google Scholar
Biggs J, Tang C, Kennedy G (2022) Teaching for quality learning at university, 5th edn. McGraw-Hill Education, UK
Boudouaia A, Mouas S, Kouider B (2024) A study on ChatGPT-4 as an innovative approach to enhancing English as a foreign language writing learning. J Educ Comput Res 07356331241247465. https://doi.org/10.3991/ijep.v13i8.40853
Bouzar A, EL Idrissir K, Ghourdou T (2024) ChatGPT and academic writing self-efficacy: unveiling correlations and technological dependency among postgraduate students. Arab World English J Special Issue on ChatGPT. https://doi.org/10.24093/awej/chatgpt.15
Chan CKY, Zhou W (2023) An expectancy value theory (EVT) based instrument for measuring student perceptions of generative AI. Smart Learn Environ 10(1):64. https://doi.org/10.1186/s40561-023-00284-4
Article Google Scholar
Chan CKY, Lee KK (2023) The AI generation gap: are Gen Z students more interested in adopting generative AI such as ChatGPT in teaching and learning than their Gen X and millennial generation teachers?. Smart Learn Environ 10(1):60. https://doi.org/10.1186/s40561-023-00269-3
Article Google Scholar
Chang CY, Yang CL, Jen HJ, Ogata H, Hwang GH (2024) Facilitating nursing and health education by incorporating ChatGPT into learning designs. Educ Technol Soc 27(1):215–230. https://doi.org/10.30191/ETS.202401_27(1).TP02
Article Google Scholar
Chen K, Yang S, Luh D, Chen Z, Ai H, An Y (2025) Gamification as an innovative approach for the assessment of procedural knowledge. Electronics 14(8):1573. https://doi.org/10.3390/electronics14081573
Article Google Scholar
Chen CH, Chang CL (2024) Effectiveness of AI-assisted game-based learning on science learning outcomes, intrinsic motivation, cognitive load, and learning behavior. Educ Inf Technol 1–22. https://doi.org/10.1007/s10639-024-12553-x
Chi MT, Wylie R (2014) The ICAP framework: Linking cognitive engagement to active learning outcomes. Educ Psychol 49(4):219–243. https://doi.org/10.1080/00461520.2014.965823
Article Google Scholar
Cohen J(1992) A power primer Psychol Bull 112:155–159. https://doi.org/10.1037/0033-2909.112.1.155
Article PubMed CAS Google Scholar
Darmawansah, D., Rachman, D., Febiyani, F., & Hwang, G. J. (2025). ChatGPT-supported collaborative argumentation: Integrating collaboration script and argument mapping to enhance EFL students’ argumentation skills. Education and information technologies, 30(3), 3803-3827. https://doi.org/10.1007/s10639-024-12986-4
Darvishi A, Khosravi H, Sadiq S, Gašević D, Siemens G (2024) Impact of AI assistance on student agency. Comput Educ 210:104967. https://doi.org/10.1016/j.compedu.2023.104967
Article Google Scholar
Dasari D, Hendriyanto A, Sahara S, Suryadi D, Muhaimin LH, Chao T, Fitriana L (2024) ChatGPT in didactical tetrahedron, does it make an exception? A case study in mathematics teaching and learning. Front Educ 8, 1295413
Deng R, Jiang M, Yu X, Lu Y, Liu S (2025) Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies. Comput Educ 227:105224. https://doi.org/10.1016/j.compedu.2024.105224
Article Google Scholar
Deng X, Yu Z (2023) A meta-analysis and systematic review of the effect of chatbot technology use in sustainable education. Sustainability 15(4):2940. https://doi.org/10.3390/su15042940
Article ADS Google Scholar
Er E, Akçapınar G, Bayazıt A, Noroozi O, Banihashem SK (2025) Assessing student perceptions and use of instructor versus AI-generated feedback. Br J Educ Technol 56(3):1074–1091. https://doi.org/10.1111/bjet.13558
Article Google Scholar
Escalante J, Pack A, Barrett A (2023) AI-generated feedback on writing: insights into efficacy and ENL student preference. Int J Educ Technol High Educ (20), 55. https://doi.org/10.1186/s41239-023-00425-2
Essel HB, Vlachopoulos D, Essuman AB, Amankwa JO (2024) ChatGPT effects on cognitive skills of undergraduate students: receiving instant responses from AI-based conversational large language models (LLMs). Comput Educ: Artif Intell 6:100198. https://doi.org/10.1016/j.caeai.2023.100198
Article Google Scholar
Fernández-Llamas C, Conde MA, Rodríguez-Lera FJ, Rodríguez-Sedano FJ, García F (2018) May I teach you? Students’ behavior when lectured by robotic vs. human teachers. Comput Hum Behav 80:460–469. https://doi.org/10.1016/j.chb.2017.09.028
Article Google Scholar
Furtak EM, Kunter M (2012) Effects of autonomy-supportive teaching on student learning and motivation. J Exp Educ 80(3):284–316. https://doi.org/10.1080/00220973.2011.573019
Article Google Scholar
Giannakos M, Azevedo R, Brusilovsky P, Cukurova M, Dimitriadis Y, Hernandez-Leo D, … Rienties B (2024) The promise and challenges of generative AI in education. Behav Inf Technol 1–27. https://doi.org/10.1080/0144929X.2024.2394886
Gu X, Hu M (2018) The learning effects of E-Schoolbag: a meta-analysis of 39 studies at home and abroad. e-Educ Res 39(5):19–25. https://doi.org/10.13811/j.cnki.eer.2018.05.003
Article Google Scholar
Gupta M, Akiri C, Aryal K, Parker E, Praharaj L (2023) From ChatGPT to threatgpt: Impact of generative AI in cybersecurity and privacy. IEEE Access. https://doi.org/10.1109/ACCESS.2023.3300381
Hakiki M, Fadli R, Samala AD, Fricticarani A, Dayurni P, Rahmadani K, Sabir A (2023) Exploring the impact of using Chat-GPT on student learning outcomes in technology learning: the comprehensive experiment. Adv Mob Learn Educ Res 3(2):859–872. https://doi.org/10.25082/amler.2023.02.013
Article Google Scholar
Haristiani N (2019) Artificial Intelligence (AI) chatbot as language learning medium: an inquiry. J Phys: Conf Ser 1387(1):012020. https://doi.org/10.1088/1742-6596/1387/1/012020
Harrison LM, Hurd E, Brinegar KM (2023) Critical race theory, books, and ChatGPT: moving from a ban culture in education to a culture of restoration. Middle Sch J 54(3):2–4. https://doi.org/10.1080/00940771.2023.2189862
Article Google Scholar
Hauk D, Gröschner A (2022) How effective is learner-controlled instruction under classroom conditions? A systematic review. Learn Motiv 80: 101850. https://doi.org/10.1016/j.lmot.2022.101850
Article Google Scholar
Higgins JP, Thompson SG, Deeks JJ, Altman DG (2003) Measuring inconsistency in meta-analyses. BMJ 327(7414):557–560. https://doi.org/10.1136/bmj.327.7414.557
Article PubMed PubMed Central Google Scholar
Huesca G, Martínez-Treviño Y, Molina-Espinosa JM, Sanromán-Calleros AR, Martínez-Román R, Cendejas-Castro EA, Bustos R (2024) Effectiveness of using ChatGPT as a tool to strengthen benefits of the flipped learning strategy. Educ Sci 14(6):660. https://doi.org/10.3390/educsci14060660
Article Google Scholar
Weidlich J, Gašević D, Drachsler H, Kirschner P (2025) ChatGPT in Education: An Effect in Search of a Cause. J Comput Assist Learn 41(5):e70105. https://doi.org/10.1111/jcal.v41.510.1111/jcal.70105
Article Google Scholar
Jeno LM, Vandvik V, Eliassen S, Grytnes JA (2019) Testing the novelty effect of an m-learning tool on internalization and achievement: a Self-Determination Theory approach. Comput Educ 128:398–413. https://doi.org/10.1016/j.compedu.2018.10.008
Article Google Scholar
Jeong H, Hmelo-Silver CE, Jo K (2019) Ten years of computer-supported collaborative learning: a meta-analysis of CSCL in STEM education during 2005–2014. Educ Res Rev 28: 100284. https://doi.org/10.1016/j.edurev.2019.100284
Article Google Scholar
Jin F, Sun L, Pan Y, Lin CH (2025) High heels, compass, Spider-Man, or drug? Metaphor analysis of generative artificial intelligence in academic writing. Comput Educ 105248. https://doi.org/10.1016/j.compedu.2025.105248
Ju Q (2023) Experimental evidence on negative impact of generative AI on scientific learning outcomes. arXiv:2311.05629
Karaman MR (2024) Are lesson plans created by ChatGPT More Effective? An experimental study. Int J Technol Educ 7(1):107–127. https://doi.org/10.46328/ijte.607
Article Google Scholar
Kavadella A, Da Silva MAD, Kaklamanos EG, Stamatopoulos V, Giannakopoulos K (2024) Evaluation of ChatGPT’s real-life implementation in undergraduate dental education: mixed methods study. JMIR Med Educ 10(1):e51344. https://doi.org/10.2196/51344
Article PubMed PubMed Central Google Scholar
Khowaja SA, Khuwaja P, Dev K, Wang W, Nkenyereye L (2024) ChatGPT needs spade (sustainability, privacy, digital divide, and ethics) evaluation: a review. Cogn Comput 16(5):2528–2550. https://doi.org/10.1007/s12559-024-10285-1
Article Google Scholar
Kim S, Lee J (2023) The effect of voice-enabled ChatGPT on speaking proficiency, affective and metacognitive domains of Korean EFL learners. J Res Curric Instr 27(4):344–359. https://doi.org/10.24231/rici.2023.27.4.344
Article Google Scholar
Kohnke L, Zou D, Su F (2025) Exploring the potential of GenAI for personalised english teaching: learners’ experiences and perceptions. Comput Educ: Artif Intell 100371. https://doi.org/10.1016/j.caeai.2025.100371
Kosar T, Ostojić D, Liu YD, Mernik M (2024) Computer science education in ChatGPT era: experiences from an experiment in a programming course for novice programmers. Mathematics 12(5):629. https://doi.org/10.3390/math12050629
Article Google Scholar
Laun M, Wolff F (2025) Chatbots in education: Hype or help? A meta-analysis. Learn Individ Differ 119: 102646. https://doi.org/10.1016/j.lindif.2025.102646
Article Google Scholar
Lazonder AW, Harmsen R (2016) Meta-analysis of inquiry-based learning: effects of guidance. Rev Educ Res 86(3):681–718. https://doi.org/10.2307/24752879
Article Google Scholar
Lee HY, Chen PH, Wang WS, Huang YM, Wu TT (2024) Empowering ChatGPT with guidance mechanism in blended learning: effect of self-regulated learning, higher-order thinking skills, and knowledge construction. Int J Educ Technol High Educ 21(1):16. https://doi.org/10.1186/s41239-024-00447-4
Article Google Scholar
Lee YF, Hwang GJ, Chen PY (2022) Impacts of an AI-based chatbot on college students’ after-class review, academic performance, self-efficacy, learning attitude, and motivation. Educ Technol Res Dev 70(5):1843–1865. https://doi.org/10.1007/s11423-022-10142-8
Article Google Scholar
Li H (2023) Effects of a ChatGPT-based flipped learning guiding approach on learners’ courseware project performances and perceptions. Australas J Educ Technol 39(5):40–58. https://doi.org/10.14742/ajet.8923
Article Google Scholar
Li T, Ji Y, Zhan Z (2024) Expert or machine? Comparing the effect of pairing student teacher with in-service teacher and ChatGPT on their critical thinking, learning performance, and cognitive load in an integrated-STEM course. Asia Pac J Educ 44(1):45–60. https://doi.org/10.1080/02188791.2024.2305163
Article ADS Google Scholar
Li H, Wang W (2023) Human–computer collaborative deep exploratory teaching mode: taking the human–computer collaborative exploratory learning system developed based on ChatGPT and QQ as an example. Open Educ Res (6), 69–81. https://doi.org/10.13966/j.cnki.kfjyyj.2023.06.008
Li H, Wang W, Li G, Wang Y (2024) Teaching method of intelligent midwifery: taking the teaching practice of Socratic Chatbot as an example. Open Educ Res (2), 89–99. https://doi.org/10.13966/j.cnki.kfjyyj.2024.02.010
Limo FAF, Tiza DRH, Roque MM, Herrera EE, Murillo JPM, Huallpa JJ, Gonzáles JLA (2023) Personalized tutoring: ChatGPT as a virtual tutor for personalized learning experiences. Przestrz Społecz (Soc Space) 23(1):293–312
Google Scholar
Liu W, Wang Y (2024) The effects of using AI tools on critical thinking in English literature classes among EFL learners: an intervention study. Eur J Educ 59(4):e12804. https://doi.org/10.1111/ejed.12804
Article Google Scholar
Liu GL, Darvin R, Ma C (2024) Exploring AI-mediated informal digital learning of English (AI-IDLE): a mixed-method investigation of Chinese EFL learners’ AI adoption and experiences. Comput Assist Language Learn 1–29. https://doi.org/10.1080/09588221.2024.2310288
Liu Z, He X, Liu L, Liu T, Zhai X Context Matters: A Strategy to Pre-train Language Model for Science Education. In Wang, N. Rebolledo-Mendez, G. Dimitrova, V. Matsuda, N. & Santos, OC. (Eds.), Artificial Intelligence in Education. Posters and Late Breaking Results, Workshops and Tutorials, Industry and Innovation Tracks, Practitioners, Doctoral Consortium and Blue Sky, Vol. 1831, pp. 666–674 (Springer, Cham. 2023) https://doi.org/10.1007/978-3-031-36336-8_103
Liu Z, Zhang W, Yang P (2025) Can AI chatbots effectively improve EFL learners’ learning effects?—A meta-analysis of empirical research from 2022–2024. Comput Assisted Language Learn 1–27. https://doi.org/10.1080/09588221.2025.2456512
Lo CK, Hew KF, Jong MSY (2024) The influence of ChatGPT on student engagement: a systematic review and future research agenda. Comput Educ 105100. https://doi.org/10.1016/j.compedu.2024.105100
Lohse KR, Healy AF (2012) Exploring the contributions of declarative and procedural information to training: a test of the procedural reinstatement principle. J Appl Res Mem Cogn 1(2):65–72. https://doi.org/10.1016/j.jarmac.2012.02.002
Article Google Scholar
Luo J (2025) How does GenAI affect trust in teacher-student relationships? Insights from students’assessment experiences. Teach High Educ 30(4):991–1006. https://doi.org/10.1080/13562517.2024.2341005
Article Google Scholar
Luo H, Liao X, Ru Q, Wang Z (2024) Generative AI-supported teacher comments: an empirical study based on junior high school mathematics classrooms. e-Educ Res (5), 58–66. https://doi.org/10.13811/j.cnki.eer.2024.05.008
Mogavi RH, Den C, Kim JJ, Zhou P, Kwon YD. Metwally, AHS. … & Hui, P. (2023). Exploring user perspectives on ChatGPT: Applications, perceptions, and implications for AI-integrated education. Computers in Human Behavior: Artificial Humans, 100027. https://doi.org/10.1016/j.chbah.2023.100027
Montuori C, Gambarota F, Altoé G, Arfé B (2023) The cognitive effects of computational thinking: a systematic review and meta-analytic study. Comput Educ 104961. https://doi.org/10.1016/j.compedu.2023.104961
Murgia E, Pera MS, Landoni M, & Huibers, T. (2023). Children on ChatGPT readability in an educational context: Myth or opportunity? In Adjunct Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization (pp. 311–316). Association for Computing Machinery (ACM). https://doi.org/10.1145/3563359.3596996
Nakavachara V, Potipiti T, & Chaiwat T. (2024) Experimenting with generative AI: does ChatGPT really increase everyone’s productivity? Preprint at arXiv preprint arXiv:2403.01770
Newton P. & Xiromeriti M. (2024). ChatGPT performance on multiple choice question examinations in higher education. A pragmatic scoping review. Assessment & Evaluation in Higher Education, 49(6), 781–798. https://doi.org/10.1080/02602938.2023.2299059
Ng DTK, Tan CW, Leung JKL (2024) Empowering student self-regulated learning and science education through ChatGPT: a pioneering pilot study. Br J Educ Technol. https://doi.org/10.1111/bjet.13454
Nguyen A, Hong Y, Dang B, Huang X (2024) Human–AI collaboration patterns in AI-assisted academic writing. Stud High Educ 1–18. https://doi.org/10.1080/03075079.2024.2323593
Othman K (2023) Towards implementing AI Mobile application chatbots for EFL learners at primary schools in Saudi Arabia. J Namib Stud: Hist Politics Cult 33:271–287. https://doi.org/10.59670/jns.v33i.434
Article Google Scholar
Page MJ, Moher D (2017) Evaluations of the uptake and impact of the preferred reporting Items for Systematic reviews and Meta-Analyses (PRISMA) statement and extensions: a scoping review. Syst Rev 6:1–14. https://doi.org/10.1186/s13643-017-0663-8
Article Google Scholar
Pallivathukal RG, Soe HHK, Donald PM, Samson RS, Ismail ARH (2024) ChatGPT for academic purposes: survey among undergraduate healthcare students in Malaysia. Cureus 16(1). https://doi.org/10.7759/cureus.53032
Pellas N (2023) The effects of generative AI platforms on undergraduates’ narrative intelligence and writing self-efficacy. Educ Sci 13(11):1155. https://doi.org/10.3390/educsci13111155
Article Google Scholar
Rees T (2022) Non-human words: on GPT-3 as a philosophical laboratory. Daedalus 151(2):168–182. https://doi.org/10.1162/daed_a_01908
Article Google Scholar
Rudolph J, Tan S, Tan S (2023) ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?. J Appl Learn Teach 6(1):342–363. https://doi.org/10.37074/jalt.2023.6.1.9
Article Google Scholar
Ryle G, Tanney J (2009) The concept of mind. Routledge
Seo K, Tang J, Roll I, Fels S, Yoon D (2021) The impact of artificial intelligence on learner–instructor interaction in online learning. Int J Educ Technol High Educ 18:1–23. https://doi.org/10.1186/s41239-021-00292-9
Article Google Scholar
Shelby LB, Vaske JJ (2008) Understanding meta-analysis: a review of the methodological literature. Leis Sci 30(2):96–110. https://doi.org/10.1080/01490400701881366
Article Google Scholar
Shoufan A (2023) Can students without prior knowledge use ChatGPT to answer test questions? An empirical study. ACM Trans Comput Educ 23(4):1–29. https://doi.org/10.1145/3628162
Article Google Scholar
Song C, Song Y (2023) Enhancing academic writing skills and motivation: assessing the efficacy of ChatGPT in AI-assisted language learning for EFL students. Front Psychol 14: 1260843. https://doi.org/10.3389/fpsyg.2023.1260843
Article PubMed PubMed Central Google Scholar
Strzelecki A (2023) To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technology. Interact Learn Environ 1–14. https://doi.org/10.1080/10494820.2023.2209881
Sun D, Boudouaia A, Zhu C, Li Y (2024) Would ChatGPT-facilitated programming mode impact college students’ programming behaviors, performances, and perceptions? An empirical study. Int J Educ Technol High Educ 21(1):14. https://doi.org/10.1186/s41239-024-00446-5
Article Google Scholar
Sun L, Zhou L (2024) Does generative artificial intelligence improve the academic achievement of college students? A meta-analysis. J Educ Comput Res. https://doi.org/10.1177/07356331241277937
Susnjak T, McIntosh TR (2024) ChatGPT: the end of online exam integrity?. Educ Sci 14(6):656. https://doi.org/10.3390/educsci14060656
Article Google Scholar
Tian L, Zhou Y (2020) Learner engagement with automated feedback, peer feedback and teacher feedback in an online EFL writing context. System 91: 102247. https://doi.org/10.1016/j.system.2020.102247
Article Google Scholar
Uddin SJ, Albert A, Tamanna M, Ovid A, Alsharef A (2024) ChatGPT as an educational resource for civil engineering students. Comput Appl Eng Educ e22747. https://doi.org/10.1002/cae.22747
Ur Rehman A, Behera RK, Islam MS, Abbasi FA, Imtiaz A (2024) Assessing the usage of ChatGPT on life satisfaction among higher education students: the moderating role of subjective health. Technol Soc 102655. https://doi.org/10.1016/j.techsoc.2024.102655
Urban M, Děchtěrenko F, Lukavský J, Hrabalová V, Svacha F, Brom C, Urban K (2024) ChatGPT improves creative problem-solving performance in university students: an experimental study. Comput Educ 215: 105031. https://doi.org/10.1016/j.compedu.2024.105031
Article Google Scholar
Urban M, Brom C, Lukavský J, Děchtěrenko F, Hein V, Svacha F, Urban K (2025) ChatGPT can make mistakes. Check important info. Epistemic beliefs and metacognitive accuracy in students’ integration of ChatGPT content into academic writing. Br J Educ Technol. https://doi.org/10.1111/bjet.13591
Vasconcelos MAR. & Santos, RPD. (2023). Enhancing STEM learning with ChatGPT and Bing Chat as objects to think with: A case study. Eurasia Journal of Mathematics, Science & Technology Education, *19*(7), em2296, 1–15. https://doi.org/10.29333/ejmste/13313
Venkatesh V, Morris MG, Davis GB, Davis FD (2003) User acceptance of information technology: toward a unified view. MIS Q 425–478. https://doi.org/10.2307/30036540
Villena-Taranilla R, Tirado-Olivares S, Cózar-Gutiérrez R, González-Calero JA (2022) Effects of virtual reality on learning outcomes in K-6 education: a meta-analysis. Educ Res Rev 35:100434. https://doi.org/10.1016/j.edurev.2022.100434
Article Google Scholar
Wang J, Fan W (2025) The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: insights from a meta-analysis. Humanit Soc Sci Commun 12(1):1–21. https://doi.org/10.1057/s41599-025-04787-y
Article Google Scholar
Wang M, Wang M, Xu X, Yang L, Cai D, Yin M (2023) Unleashing ChatGPT’s power: a case study on optimizing information retrieval in flipped classrooms via prompt engineering. IEEE Trans Learn Technol. https://doi.org/10.1109/TLT.2023.3324714
Wang Y, Wang L, Siau KL (2024) Human-centered interaction in virtual worlds: a new era of generative artificial intelligence and metaverse. Int J Hum–Comput Interact 1–43. https://doi.org/10.1080/10447318.2024.2316376
Wang Z, Ma Y, Yang X, Li K (2023) Impact of ChatGPT-based reading platforms on the academic reading ability of graduate students. Open Educ Res (6), 60–68. https://doi.org/10.13966/j.cnki.kfjyyj.2023.06.007
Widianingtyas N, Mukti TWP, Silalahi RMP (2023) ChatGPT in language education: perceptions of teachers—a beneficial tool or potential threat?. Voices Engl Lang Educ Soc 7(2):279–290. https://doi.org/10.29408/veles.v7i2.20326
Article Google Scholar
Wilson DB, Lipsey MW (2001) The role of method in treatment effectiveness research: evidence from meta-analysis. Psychol Methods 6(4):413. https://doi.org/10.1037/1082-989x.6.4.413
Article PubMed CAS Google Scholar
Wu R, Yu Z (2024) Do AI chatbots improve students learning outcomes? Evidence from a meta-analysis. Br J Educ Technol 55(1):10–33. https://doi.org/10.1111/bjet.13334
Article MathSciNet Google Scholar
Wu TT, Lee HY, Li PH, Huang CN, Huang YM (2024) Promoting self-regulation progress and knowledge construction in blended learning via ChatGPT-based learning aid. J Educ Comput Res 61(8):3–31. https://doi.org/10.1177/07356331231191125
Article Google Scholar
Wu X, Li R (2024) Effects of robot-assisted language learning on English-as-a-foreign-language skill development. J Educ Comput Res 62(4):1010–1034. https://doi.org/10.1177/07356331231226171
Article Google Scholar
Wu X, Yang Y, Zhou X, Xia Y, Liao H (2024b) A meta-analysis of interdisciplinary teaching abilities among elementary and secondary school STEM teachers. Int J STEM Educ 11(1):38. https://doi.org/10.1186/s40594-024-00500-8
Article Google Scholar
Wu JH, Zhou WT, Cao C (2024) An empirical study on empowering oral teaching with AIGC. China Educ Technol (4), 105–111. https://doi.org/10.3969/j.issn.1006-9860.2024.04.014
Wu X, Liao H, Guan L (2024a) Examining the influencing factors of elementary and high school STEM teachers’self-efficacy: a meta-analysis. Curr Psychol 2024:1–17.https://doi.org/10.1007/s12144-024-06227-7
Wu XN (2023) STEM course development: theory and practice. Guangming Daily Publishing House
Xing W, Song Y, Li C, Liu Z, Zhu W, Oh H (2025) Development of a generative AI-powered teachable agent for middle school mathematics learning: a design-based research study. Br J Educ Technol. https://doi.org/10.1111/bjet.13586
Yilmaz R, Yilmaz FGK (2023) The effect of generative artificial intelligence (AI)-based tool use on students’ computational thinking skills, programming self-efficacy and motivation. Comput Educ: Artif Intell 4: 100147. https://doi.org/10.1016/j.caeai.2023.100147
Article Google Scholar
Yusuf A, Pervin N, Román-González M (2024) Generative AI and the future of higher education: a threat to academic integrity or reformation? Evidence from multicultural perspectives. Int J Educ Technol High Educ 21: 21. https://doi.org/10.1186/s41239-024-00453-6
Article Google Scholar
Yuxian J (2025) Bridging the knowledge-skill gap: the role of large language model and critical thinking in education. Comput Educ 105357. https://doi.org/10.1016/j.compedu.2025.105357
Zhai X, Chu X, Jiao L, Tong Z, Li Y (2023) An empirical study on the effectiveness of human–computer collaborative learning based on generative artificial intelligence + metaverse. Open Educ Res (5), 26–36. https://doi.org/10.13966/j.cnki.kfjyyj.2023.05.003
Zheng L, Gao L, Huang Z (2024) Can chatbots based on generative artificial intelligence facilitate online collaborative learning performance? e-Educ Res (3), 70–76+84. https://doi.org/10.13811/j.cnki.eer.2024.03.010
Zhou R, He X, Fan Q, Li Y, Li Y, Xiao X, Fang J (2025) Exploring ChatGPT-facilitated scaffolding in undergraduates’ mathematical problem solving. J Comput Assist Learn 41(4):e70077. https://doi.org/10.1111/jcal.70077
Article Google Scholar
Zhou W, Kim Y (2024) Innovative music education: an empirical assessment of ChatGPT-4’s impact on student learning experiences. Educ Inf Technol 1–27. https://doi.org/10.1007/s10639-024-12705-z
Zhu G, Fan X, Hou C, Zhong T, Seow P, Shen-Hsing AC, Poh TL (2023) Embrace opportunities and face challenges: using ChatGPT in undergraduate students’ collaborative interdisciplinary learning. Preprint at https://doi.org/10.48550/arXiv.2305.18616
Zhu Y, Liu Q, Zhao L (2025) Exploring the impact of generative artificial intelligence on students’ learning outcomes: a meta-analysis. Educ Inf Technol 1–29. https://doi.org/10.1007/s10639-025-13420-z
Download references
This research was funded by the General Project of the National Social Science Fund (Education) (Grant No. BIA250124), and a project supported by the Scientific Research Fund of Hunan Provincial Education Department (Grant No. 25A0349).
School of Education, Hunan University of Science and Technology, Xiangtan, China
Xinning Wu, Pei Zhu, Jinliang Zhang, Mengwei Yin & Yingxi Wang
PubMed Google Scholar
PubMed Google Scholar
PubMed Google Scholar
PubMed Google Scholar
PubMed Google Scholar
Conceptualisation: XW; methodology: PZ; software: MY; validation: MY; data curation: PZ; writing—original draft preparation: XW and PZ; writing—review and editing: XW and JZ; visualisation: MY; supervision: JZ and WY. All authors have read and agreed to the published version of the manuscript.
Correspondence to Xinning Wu.
The authors declare no competing interests.
This study does not involve human participants; their personal or biological data and ethical approval to use them were unnecessary since the authors developed a meta-analysis by working with empirical studies.
This study does not involve human participants, so their consent was not required. The nature of this study (a meta-analysis) did not require informing any human participant.
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Reprints and permissions
Wu, X., Zhu, P., Zhang, J. et al. ChatGPT’s impact on student learning outcomes: a meta-analysis of 35 experimental studies. Humanit Soc Sci Commun 13, 684 (2026). https://doi.org/10.1057/s41599-026-07019-z
Download citation
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1057/s41599-026-07019-z
Anyone you share the following link with will be able to read this content:
Sorry, a shareable link is not currently available for this article.
Provided by the Springer Nature SharedIt content-sharing initiative
Advertisement
Humanities and Social Sciences Communications (Humanit Soc Sci Commun)
ISSN 2662-9992 (online)
© 2026 Springer Nature Limited