People accepted 78% of wrong AI actions when uncertainty stayed hidden
New analysis today of a 30 September preprint: in a controlled puzzle study, people often approved incorrect AI moves when the system did not reveal ambiguity. Targeted warnings helped, but the best-performing warning relied on oracle knowledge that a real product would not have.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Can a deployable uncertainty signal reduce mistaken human acceptance without creating so many false warnings that people ignore or overrule correct AI actions?
At a glance
- 1The evidence is an unreviewed preprint submitted on 30 September; this is new analysis today, not a claim of a 1 October source.
- 2In a 210-person UK study, participants given the AI's default message accepted 78% of wrong placements and could not distinguish right from wrong better than chance.
- 3An oracle-targeted hedge reduced wrong-move acceptance to 36%, but the model's own uncertainty signal was too weak to deliver the same protection reliably.
Living evidence record
Impact record IAI-0VC7H5B
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
1 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what arXiv / Microsoft Research and Harvard University published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
New analysis today — the source was submitted on 30 September
A new preprint examines a practical failure in human–AI teamwork: an AI can be unsure which object a person means, yet act as though the instruction were clear. The paper was submitted to arXiv at 11:17 UTC on 30 September. This report is new analysis today, not a claim that the source was published on 1 October. The study has not been peer reviewed or independently replicated, and this record relies on one primary research source.
The authors created a collaborative puzzle in which a human Helper describes one of 24 visually ambiguous pieces and an AI Worker chooses where to place it. The central question is not whether a model can produce a fluent reply. It is whether the model recognises competing possible referents, communicates uncertainty and gives the human enough information to stop a wrong action.[1]
What the researchers measured
The model experiments tested GPT-4.1, GPT-5 and GPT-5.5 in text and vision-language versions of the task. The researchers compared two uncertainty signals. One was the probability attached to the model's action tokens. The other came from separately asking the model to distribute belief across candidate pieces. They evaluated calibration with expected calibration error, or ECE, and tested whether confidence distinguished correct from incorrect placements with area under the receiver operating characteristic curve, or AUROC.
The human experiment recruited 210 English-fluent adults in the United Kingdom through Prolific, aged 18 to 76. Between 41 and 43 people were assigned to each of five conditions. Participants judged recorded puzzle turns under different combinations of message precision, visibility and hedging. The paper reports 107 men, 101 women and two participants who did not disclose gender; compensation was £12.50 per hour and the study received institutional ethics approval.[1]
The models were confident when they were wrong
Raw action-token probabilities were severely overconfident: mean confidence was 0.97 while placement accuracy was 53%, producing an ECE of 0.44. The separately elicited belief distribution was better calibrated, with ECE 0.15, and better at separating correct from incorrect placements, with AUROC 0.65. Those values show a useful signal, not a dependable warning system; an AUROC of 0.65 remains far from perfect discrimination.
Uncertainty rose when a person's description became vaguer, but barely responded when visually confusable alternative pieces appeared—even though those alternatives increased errors. The models seldom asked for clarification. Depending on model and condition, clarification occurred on 3.5% to 16.7% of image turns; the highest text-only rate reported was 23.8% for GPT-5. The study therefore identifies a mismatch between internal uncertainty, contextual ambiguity and what the human partner is told.[1]
People accepted 78% of wrong moves when the warning was missing
In the generic condition where participants could not see the Worker's board and received its default message, they accepted 78% of wrong placements. Their ability to distinguish correct from incorrect moves was no better than chance, with an AUC of 0.50. This is a controlled judgement task rather than a live workplace trial, but it demonstrates how a polished message can hide the information a person needs to intervene.
More precise descriptions and a targeted hedge reduced acceptance of wrong moves to 36% while preserving 74% acceptance of correct moves. The strongest hedge was generated with knowledge of whether the model's chosen piece was actually correct, so it is an oracle condition rather than a deployable product. A hedge driven by the model's own uncertainty signal had only about 0.55 AUROC and could make decisions worse when it warned at the wrong times.[1]
Impact on people and global context
For people using AI assistants, the immediate risk is misplaced trust. In customer service, clinical administration, industrial maintenance or office automation, a model may confidently act on the wrong record, image or instruction while giving the operator no reason to pause. A useful interface should expose the alternatives the system considered, ask a focused clarification question and delay consequential action when the intended referent is uncertain.
The participants were UK-based English speakers and the task used controlled visual puzzles, so the reported percentages should not be generalised to other languages, cultures or professional settings. Ambiguity may be harder across dialects, translation and specialised terminology. The author group spans Microsoft Research UK and Harvard University in the United States, but the study does not compare countries or evaluate population-level outcomes.[1]
Evidence, limitations and what to watch
The experiment measured participant judgements and intended acceptance of recorded moves, not behaviour in an active collaboration with real consequences. It tested three models from one provider, a single puzzle design and selected early trials. The preprint has no independent replication, and the accessible manuscript does not provide a funding or competing-interest statement. Microsoft Research affiliations create a commercial institutional interest that readers should keep visible without treating it as evidence that the results are invalid.
The next test is whether uncertainty prompts generated without oracle knowledge reduce real errors while avoiding excessive warnings. Replications should include other model families, languages, accessibility needs and high-consequence workflows, and should measure completed actions rather than stated intentions. Until then, the defensible conclusion is narrow: people cannot reliably correct an AI's mistaken interpretation when the system does not communicate referential uncertainty accurately.[1]
What this means for people
- A fluent AI response can conceal ambiguity, leaving people unable to recognise when the system has acted on the wrong object, record or instruction.
- Interfaces should ask focused clarification questions and delay consequential actions when several plausible referents remain.
- Warnings can also cause harm when poorly targeted, so products need evidence that uncertainty prompts improve decisions rather than merely adding cautionary language.
Global context
The human sample consisted of UK-based English speakers, while the authors are affiliated with Microsoft Research UK and Harvard University. The experiment does not establish outcomes in other languages, countries or professional settings; ambiguity may change with dialect, translation, specialist vocabulary and local interaction norms.
What the evidence does not yet show
- This is a single unreviewed preprint with no independent replication and no second source validating the numerical results.
- The 210-person study measured judgements about recorded puzzle turns, not completed actions in a live, consequential workflow.
- Only three OpenAI model generations and one task design were tested, so the results do not estimate failure rates for other providers or deployed products.
- The strongest hedge used oracle knowledge of whether the action was correct; the accessible manuscript provides no funding or competing-interest statement.
What to watch next
- Independent replications using other model families, languages, task types and real collaborative decisions.
- Whether non-oracle uncertainty signals can improve human intervention without generating excessive false alarms.
- Product interfaces that expose alternative interpretations and require clarification before consequential action.
Evidence trail
Sources used for this report
Links checked 1 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
AI agents often abandoned the plans they declared; stronger execution structure helped in benchmarks
New analysis today of a 29 September preprint led from Abu Dhabi: generic agents often failed to preserve their stated reasoning pattern. Structured executors improved some benchmark results, but the study is unreviewed and its fidelity labels received only limited human validation.
6 min · 2 sources
Science & Research
Collective intelligence becomes a design principle for AI-assisted chemical synthesis
A Nature study examines how human expertise and machine systems can be combined to plan chemical synthesis, moving beyond a single-model recommendation toward a collective workflow.
4 min · 1 source
Science & Research
Can AI design better optimisation rules?
A Nature Machine Intelligence paper reports that a structured LLM framework outperformed prior LLM approaches across 36 combinatorial-optimisation benchmarks and produced feasible algorithms for four new port-logistics problems. It remains an offline benchmark, not a live operations result.
8 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.