Back to the news portal
EducationResearch paperResearchSource analysisAsiaChina

Did ChatGPT improve medical training—or the whole teaching package?

A peer-reviewed randomised study at a Chinese hospital found higher examination scores after combining problem-based learning with ChatGPT. Because there was no PBL-only arm, the trial cannot isolate the AI's contribution.

By The Impact of AI Research DeskReleased 2 October 2026 at 02:07 BST8 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesMedical educationChatGPTProblem-based learningClinical trainingEducation research

Research topic

Whether a combined problem-based-learning and ChatGPT package improves laboratory-medicine film-reading education

The Impact of AI research cover asking whether ChatGPT or the whole teaching package improved medical training, with conceptual trainees discussing laboratory slides and a generic AI dialogue panel.
AI-generated editorial illustration. The trainees, dialogue panel and laboratory slides are conceptual and do not depict participants, patient material, an actual diagnosis or study measurements.

At a glance

  • 1Eighty residents at Xiangya Hospital were randomised 1:1 to a combined PBL-ChatGPT package or traditional lecture-based teaching, with allocation concealed in sequential opaque envelopes.
  • 2Median theory, film-reading and combined scores were higher in the intervention arm, with large rank-biserial effect sizes, but no baseline film-reading test was conducted.
  • 3The two-arm design cannot distinguish the contribution of ChatGPT from problem-based learning, and assessments were short-term, single-centre and scored by the unblinded teaching team.

Living evidence record

Impact record IAI-1YWSZX4

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

2 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

The intervention was a teaching package, not an AI-only test

Film-reading in laboratory medicine asks trainees to connect visual patterns with clinical history, test results, diagnostic possibilities and professional terminology. The study at Xiangya Hospital, Central South University, compared traditional lecture-based instruction with a package that combined problem-based learning, case preparation, structured questioning and ChatGPT 4.0. The headline question is therefore not whether a chatbot independently taught residents better. It is whether the whole redesigned learning experience outperformed the existing approach.

Residents in the intervention arm received an orientation on prompts and interpreting outputs, then completed a classroom test of their basic ChatGPT skills. Cases were sent a week in advance. Trainees reviewed laboratory information, discussed cases with the model and prepared questions. Instructors could use ChatGPT to generate layered prompts or anticipate misconceptions, but they remained responsible for checking outputs, guiding discussion, correcting descriptions and linking images to clinical reasoning.[1][2]

Eighty residents were randomised, with concealed allocation

The trial enrolled 80 people between September 2024 and February 2026: eight dedicated laboratory-medicine residents and 72 medical residents rotating through the department. A preliminary effect-size assumption of 0.65 produced a target of at least 38 participants per arm; the researchers enrolled 40 in each. Assignment used a random-number table, and an independent staff member managed sequentially numbered, opaque, sealed envelopes to conceal the next allocation.

The groups were similar in mean age, gender, previous training duration and reported prior AI use. The intervention group had a mean age of 24.2 years and the comparison group 24.1. Prior AI use was reported by 15 of 40 intervention participants and 12 of 40 controls. Both groups used unified teaching syllabi and case databases, and instructors received standardised preparation. Those controls reduce some differences, but the paper does not report a baseline film-reading proficiency test.[2]

The comparison was active discussion versus conventional instruction

Traditional classes used image demonstrations and lectures, with the teacher as the main transmitter of knowledge. The combined arm reorganised the sequence into preparation, implementation and evaluation. Residents arrived having worked through a case; instructors posed progressive questions; participants described morphology, proposed diagnostic plans and compared alternatives. ChatGPT was an aid inside that process, not the examiner or autonomous clinical decision-maker.

This difference is educationally meaningful but analytically difficult. Problem-based learning already changes preparation, participation and feedback. The intervention also gave trainees prompt instruction and a week of structured advance work. Any advantage could come from PBL, the chatbot, greater preparation time, novelty, more active facilitation or their interaction. A three-arm design—traditional teaching, PBL without ChatGPT and PBL with ChatGPT—would be needed to estimate the incremental AI contribution.[2]

Test scores favoured the combined package

After the teaching sessions, each group completed a theory examination and a film-reading examination, each scored out of 100. Median theory scores were 89 in the combined arm versus 82 with traditional teaching, a median difference of seven points and a rank-biserial correlation of 0.72. Median film-reading scores were 94 versus 90, a four-point difference with an effect size of 0.65. The combined total was 184 versus 172, with an effect size of 0.81. Reported p values were below 0.001.

A standardised 21-item film-reading assessment also favoured the intervention overall: median scores were 95.0 versus 90.5, with a five-point estimated difference and a 95% confidence interval from three to seven. Preparation, teaching methodology and overall-evaluation domains were higher after correction for four comparisons. The process domain did not differ significantly. Sensitivity analysis excluding the eight dedicated laboratory residents did not change the main examination findings.[2]

The assessment design may inflate the apparent advantage

The instructors and residents could not be blinded to the teaching method. More importantly, the in-house examinations were scored by the same teaching team that delivered the intervention. That creates room for expectancy, assessment and social-desirability bias even when rubrics are used. The study did not use an external blinded examiner, an objective structured clinical examination or performance on unseen cases at another institution.

The absence of baseline proficiency means randomisation is doing substantial work. With only 40 participants per arm, an unmeasured imbalance in prior film-reading skill could still matter. The accepted manuscript also contains an apparent internal inconsistency in its baseline table: the prose describes four laboratory and 36 rotating residents in each arm, while one table entry displays a different control-group split even as it reports identical distributions. The overall denominator is clear, but the discrepancy should be corrected in the final Version of Record.[2]

Short-term examination gains are not yet clinical competence

The study evaluates immediate educational outcomes and participant satisfaction. It does not show that knowledge was retained, that residents performed better months later, or that they made safer diagnostic decisions in clinical practice. No patient outcome, diagnostic-error rate or workload outcome was measured. The film-reading assessment includes items about the quality of the teaching session as well as individual performance, so it is not a pure measure of each trainee's skill.

The result is still useful. Random assignment and concealed allocation make it more informative than a voluntary classroom survey, and the reported differences are not trivial within the study's own scales. But the safest interpretation is that the combined package produced better short-term scores in this setting. It does not establish that ChatGPT alone works, that the effect generalises to other specialties or countries, or that higher scores translate into better care.[2]

Teachers remain the safety layer

The paper acknowledges that model outputs can be biased, incomplete, outdated or misleading. In the tested package, physicians selected cases, verified material, corrected trainees and kept the model inside supervised discussion. That is a markedly different risk profile from allowing a resident to use an unsupervised chatbot with identifiable patient information. Any implementation should use approved tools, remove or protect confidential data and make clear that generated explanations require expert review.

For educators, the practical choice is not simply to buy access to a chatbot. The intervention required case preparation, instructor training, prompt orientation, structured questioning and active feedback. Those inputs have costs. Institutions should record preparation time, licence and infrastructure costs, access inequities and the burden of checking outputs. A teaching package that raises scores but consumes substantially more staff time may still be worthwhile, but that trade-off needs measurement rather than assumption.[1][2]

What would change the assessment

A multicentre three-arm trial would provide the clearest next step. It should test traditional instruction against PBL alone and the same PBL programme with ChatGPT, use baseline and follow-up assessments, and have blinded external examiners score unseen cases. Pre-registration, a published analysis plan and open de-identified outcomes would reduce flexibility. Longer follow-up could show whether gains persist and whether trainees become over-reliant on suggestions.

The study was supported by Hunan provincial research grants; the authors declared no competing interests, and the participant-level data are not public because of privacy and ethics restrictions. The evidence currently supports cautious experimentation with a supervised combined package, not replacement of instructors or claims that ChatGPT improves clinical competence. The assessment would become stronger if the AI-specific increment, durability, cost, fairness and effect on real diagnostic performance were demonstrated independently.[2]

What this means for people

  • Medical trainees may gain more active practice, but need protected learning environments and clear verification habits.
  • Teachers remain responsible for correcting model errors and ensuring that discussion does not expose patient information.
  • Patients would benefit only if educational gains persist and improve clinical decisions; this trial did not measure that step.

Global context

This was a single-centre trial in China using a national film-reading assessment framework and one specific teaching context. Medical education systems, approved AI tools, privacy rules and supervisory capacity differ internationally. The finding is a promising local result for a supervised package, not a global estimate of ChatGPT's educational effect.

What the evidence does not yet show

  • There was no PBL-only arm, so the study cannot isolate ChatGPT's contribution from the wider teaching redesign.
  • No baseline film-reading proficiency assessment was conducted.
  • The teaching team was unblinded and also scored the in-house assessments.
  • The 80-person, single-centre sample and short-term outcomes limit generalisation and say nothing about patient outcomes.
  • The accepted manuscript contains an apparent inconsistency in one baseline resident-type table entry that should be clarified.

What to watch next

  • Multicentre, three-arm trials separating PBL from the incremental effect of ChatGPT.
  • Blinded external assessment, baseline testing and retention at six or 12 months.
  • Measures of diagnostic accuracy, over-reliance, privacy, cost and instructor workload.
  • Correction of the baseline table discrepancy in the final Version of Record.

Evidence trail

Sources used for this report

Links checked 2 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Education

Can AI avoid overreacting to one wrong answer?

A diffusion-based knowledge-tracing model led four educational benchmarks, including difficult response reversals. The evaluation reuses folds for early stopping and scoring and does not test classroom decisions.

8 min · 2 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.