Can an inaccurate AI summary change what you remember seeing?
In a US online experiment with 328 analysed participants, accurate recall of a traffic sign fell from 83.6% after a consistent summary to 44.8% after a misleading one. The study isolates a classic misinformation effect; it does not measure real police reports or prove that ordinary model errors cause the same size of harm.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Whether misleading information in an AI-generated description of a witnessed event can reduce later memory accuracy
At a glance
- 1Researchers analysed 20 summaries of two short videos from ChatGPT-5.5 and Gemini 2.5 Flash-Lite; every summary contained at least seven coded errors.
- 2In a separate experiment, 328 US adults were analysed across consistent versus misleading content and AI versus human labels.
- 3Recall of the manipulated traffic sign was 83.6% after consistent information and 44.8% after misleading information; belief that the text came from AI rather than a human did not significantly change the effect.
Living evidence record
Impact record IAI-0W6UNFJ
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
1 October 2026
Source trail
2 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
The study joined a model audit to a controlled memory experiment
The researchers asked two linked questions. First, what kinds of errors appear when multimodal language models summarise simple videos? Second, if a later summary includes a false detail, can that detail alter what a person remembers? The design connects AI evaluation with the long-established psychological 'misinformation effect', in which post-event wording changes recall. It is more informative than a model benchmark alone because it measures a human outcome, while remaining a deliberately controlled simulation rather than a field study.
For the model audit, ChatGPT-5.5 and Gemini 2.5 Flash-Lite each summarised two 25-second animated car-pedestrian accident videos five times, producing 20 texts. Three researchers defined 19 central event details; two then coded omissions, critical inaccuracies, non-critical inaccuracies and invented additions, resolving disagreements by consensus. The prompt explicitly asked for a neutral, factual account of at least 300 words, so the models were not encouraged to speculate.[1][2]
The small video audit found errors in every output
All 20 summaries contained mistakes, from a minimum of seven to a maximum of 21 in texts averaging 305 words. Every summary omitted at least one central detail; across the predefined detail list, 51.6% of central details were omitted and 6.1% were inaccurately described. Nineteen of the 20 summaries omitted the collision itself, and 60% added at least one detail not present in the video. Gemini and ChatGPT showed different descriptive patterns, but the experiment was not powered as a general model ranking.
That denominator must stay visible. Two simple animated videos, one prompt and ten outputs per model cannot estimate the error rate for police footage, medical records or everyday summarisation. The authors frame the audit as a close analysis of materials used in the memory experiment, not a systematic survey of models. Its value is to show plausible error categories and motivate the human test, supported by a larger literature; it does not establish that most AI summaries omit half of all important facts.[1][2]
A single misleading detail nearly halved accurate recall
The human study recruited 360 US adults through Prolific. After exclusions and attrition, 328 remained: 82 in an AI-labelled consistent condition, 80 AI-labelled misleading, 83 human-labelled consistent and 83 human-labelled misleading. Participants first watched one version of the accident, containing either a stop or yield sign. After 24 to 48 hours they read the same 21-sentence ChatGPT-generated summary, with the sign description manipulated to be consistent or inconsistent with what they had seen, then completed a 12-item recognition test.
On the critical traffic-sign question, 83.6% of people who read consistent information answered correctly, compared with 44.8% after misleading information. The association was statistically significant (chi-squared 53.937, p<0.001). Other questions help rule out total forgetting: participants correctly identified day versus night 96.6% of the time and averaged 91.4% across 11 non-critical questions. The result is therefore specific to the manipulated detail rather than evidence that participants remembered nothing about the event.[1][2]
Calling the same text 'AI' did not protect participants
Half of participants were told the summary was produced by AI software and half that it was written by a human, although everyone read the same ChatGPT-generated text. Within both consistent and misleading conditions, the source label did not significantly alter the chance of answering the sign question correctly. Nor did self-reported trust in AI or frequency of chatbot use significantly moderate the effect. The study supports caution about relying on scepticism or familiarity as a sufficient defence.
It does not prove that every person is equally vulnerable or that labelling never helps in other contexts. The experiment manipulated one detail in one high-quality selected summary; a researcher had repeatedly prompted ChatGPT and made minor edits before selecting the text. Participants may also have doubted the claimed authorship. The authors acknowledge that the controlled design trades realism for causal clarity. Warnings, source access and forced comparison with original footage were not tested as interventions.[1][2]
The practical risk is strongest where summaries become records
The immediate concern is not that an AI creates a false belief from nothing. It is that a person who witnessed an event may later read a fluent automated account and unknowingly incorporate a wrong detail. In policing, insurance, healthcare or incident response, the human reviewer is often treated as the safeguard. If the draft itself shifts memory before comparison with primary evidence, asking that reviewer to 'check the AI' may provide less protection than policy assumes.
A safer workflow preserves the original audio, video or notes; clearly separates generated text from evidence; surfaces uncertainty and omissions; and requires verification before an AI draft becomes the durable record. The research was approved by institutional review boards and participants were debriefed. Funding came from a Georgetown computer-science chair and US National Science Foundation award 2205171. No commercial model company is listed as a funder in the paper.[1][2]
What would change the assessment
Confidence in generality would rise with preregistered replication using real-world, non-graphic material; audio-only and body-camera-like inputs; several languages and models; longer delays; and professional participants who routinely review records. Studies should compare naturally occurring model mistakes with researcher-inserted errors, report correction after re-viewing primary evidence and test whether provenance links, discrepancy warnings or independent verification reduce false recall without creating new bias.
The assessment would weaken if larger studies find the effect depends on this particular traffic-sign paradigm or disappears when reviewers follow realistic checking procedures. It would strengthen materially if field evaluations show altered reports, decisions or testimony. For now, the causal claim is narrow but important: in this controlled US sample, one misleading detail in an AI-generated summary substantially reduced accurate recall of the matching visual detail. That is enough to challenge human oversight designs that treat memory as an untouched source of truth.[1][2]
What this means for people
- Witnesses and frontline staff may become less reliable reviewers if an automated draft changes their memory before checking.
- People described in police, insurance or medical summaries could be affected by a wrong detail that becomes part of the durable record.
- Organisations need workflows that protect access to primary evidence and do not treat a human signature as proof of independent recall.
Global context
The experiment was conducted with US adults in English and uses a traffic-sign memory paradigm. Legal standards, police practice, recordkeeping and AI deployment differ internationally. The mechanism may travel, but its magnitude and practical consequences require local study—especially where reviewers lack access to original evidence or translated summaries introduce another source of distortion.
What the evidence does not yet show
- The model audit used two short animated videos, one prompt and 20 total summaries; it cannot estimate broad real-world error prevalence.
- The memory experiment manipulated one traffic-sign detail in a selected and lightly edited summary, not a naturally occurring deployed report.
- Participants were US adults recruited online through Prolific, so results may not generalise to professional reviewers or other countries.
- The paper was posted on 23 September and is analysed with that original date rather than being presented as current news.
What to watch next
- Independent preregistered replication across models, media types, languages and professional settings.
- Tests of provenance links, warnings and mandatory comparison with original evidence before a summary is accepted.
- Whether naturally generated model errors alter final reports, testimony, care decisions or incident records.
- Longer-term measures of correction: whether re-viewing primary evidence restores accurate memory.
Evidence trail
Sources used for this report
Links checked 1 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Society & Media
Can image-aware AI catch fake news before it spreads?
A new peer-reviewed model improved two fixed benchmark tests by combining article text, images and an AI-generated image description. It did not verify claims or face a live, changing news stream.
7 min · 2 sources
Society & Media
Will patients avoid doctors who say they use AI?
A preregistered experiment with 1,030 US adults found lower ratings and appointment intentions for fictional family doctors who said they used AI. The effect is relevant to patient trust, but an advert-based intention is not a real healthcare choice.
10 min · 2 sources
Society & Media
Did AI-written petitions persuade more people to act?
A peer-reviewed natural experiment covering 1.5 million Change.org petitions found that access to an embedded AI writer changed language and increased similarity, but did not improve the engagement outcomes the researchers measured.
5 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.