Back to the news portal
Science & ResearchResearch paperResearchSource analysisUnited KingdomUnited StatesInternational

People accepted 78% of wrong AI actions when uncertainty stayed hidden

New analysis today of a 30 September preprint: in a controlled puzzle study, people often approved incorrect AI moves when the system did not reveal ambiguity. Targeted warnings helped, but the best-performing warning relied on oracle knowledge that a real product would not have.

By The Impact of AI Editorial DeskReleased 1 October 2026 at 06:30 BST6 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesHuman–AI collaborationUncertaintyCalibrationHuman oversightInterface design

Research topic

Can a deployable uncertainty signal reduce mistaken human acceptance without creating so many false warnings that people ignore or overrule correct AI actions?

At a glance

  • 1The evidence is an unreviewed preprint submitted on 30 September; this is new analysis today, not a claim of a 1 October source.
  • 2In a 210-person UK study, participants given the AI's default message accepted 78% of wrong placements and could not distinguish right from wrong better than chance.
  • 3An oracle-targeted hedge reduced wrong-move acceptance to 36%, but the model's own uncertainty signal was too weak to deliver the same protection reliably.

Living evidence record

Impact record IAI-0VC7H5B

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

1 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what arXiv / Microsoft Research and Harvard University published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

New analysis today — the source was submitted on 30 September

A new preprint examines a practical failure in human–AI teamwork: an AI can be unsure which object a person means, yet act as though the instruction were clear. The paper was submitted to arXiv at 11:17 UTC on 30 September. This report is new analysis today, not a claim that the source was published on 1 October. The study has not been peer reviewed or independently replicated, and this record relies on one primary research source.

The authors created a collaborative puzzle in which a human Helper describes one of 24 visually ambiguous pieces and an AI Worker chooses where to place it. The central question is not whether a model can produce a fluent reply. It is whether the model recognises competing possible referents, communicates uncertainty and gives the human enough information to stop a wrong action.[1]

What the researchers measured

The model experiments tested GPT-4.1, GPT-5 and GPT-5.5 in text and vision-language versions of the task. The researchers compared two uncertainty signals. One was the probability attached to the model's action tokens. The other came from separately asking the model to distribute belief across candidate pieces. They evaluated calibration with expected calibration error, or ECE, and tested whether confidence distinguished correct from incorrect placements with area under the receiver operating characteristic curve, or AUROC.

The human experiment recruited 210 English-fluent adults in the United Kingdom through Prolific, aged 18 to 76. Between 41 and 43 people were assigned to each of five conditions. Participants judged recorded puzzle turns under different combinations of message precision, visibility and hedging. The paper reports 107 men, 101 women and two participants who did not disclose gender; compensation was £12.50 per hour and the study received institutional ethics approval.[1]

The models were confident when they were wrong

Raw action-token probabilities were severely overconfident: mean confidence was 0.97 while placement accuracy was 53%, producing an ECE of 0.44. The separately elicited belief distribution was better calibrated, with ECE 0.15, and better at separating correct from incorrect placements, with AUROC 0.65. Those values show a useful signal, not a dependable warning system; an AUROC of 0.65 remains far from perfect discrimination.

Uncertainty rose when a person's description became vaguer, but barely responded when visually confusable alternative pieces appeared—even though those alternatives increased errors. The models seldom asked for clarification. Depending on model and condition, clarification occurred on 3.5% to 16.7% of image turns; the highest text-only rate reported was 23.8% for GPT-5. The study therefore identifies a mismatch between internal uncertainty, contextual ambiguity and what the human partner is told.[1]

People accepted 78% of wrong moves when the warning was missing

In the generic condition where participants could not see the Worker's board and received its default message, they accepted 78% of wrong placements. Their ability to distinguish correct from incorrect moves was no better than chance, with an AUC of 0.50. This is a controlled judgement task rather than a live workplace trial, but it demonstrates how a polished message can hide the information a person needs to intervene.

More precise descriptions and a targeted hedge reduced acceptance of wrong moves to 36% while preserving 74% acceptance of correct moves. The strongest hedge was generated with knowledge of whether the model's chosen piece was actually correct, so it is an oracle condition rather than a deployable product. A hedge driven by the model's own uncertainty signal had only about 0.55 AUROC and could make decisions worse when it warned at the wrong times.[1]

Impact on people and global context

For people using AI assistants, the immediate risk is misplaced trust. In customer service, clinical administration, industrial maintenance or office automation, a model may confidently act on the wrong record, image or instruction while giving the operator no reason to pause. A useful interface should expose the alternatives the system considered, ask a focused clarification question and delay consequential action when the intended referent is uncertain.

The participants were UK-based English speakers and the task used controlled visual puzzles, so the reported percentages should not be generalised to other languages, cultures or professional settings. Ambiguity may be harder across dialects, translation and specialised terminology. The author group spans Microsoft Research UK and Harvard University in the United States, but the study does not compare countries or evaluate population-level outcomes.[1]

Evidence, limitations and what to watch

The experiment measured participant judgements and intended acceptance of recorded moves, not behaviour in an active collaboration with real consequences. It tested three models from one provider, a single puzzle design and selected early trials. The preprint has no independent replication, and the accessible manuscript does not provide a funding or competing-interest statement. Microsoft Research affiliations create a commercial institutional interest that readers should keep visible without treating it as evidence that the results are invalid.

The next test is whether uncertainty prompts generated without oracle knowledge reduce real errors while avoiding excessive warnings. Replications should include other model families, languages, accessibility needs and high-consequence workflows, and should measure completed actions rather than stated intentions. Until then, the defensible conclusion is narrow: people cannot reliably correct an AI's mistaken interpretation when the system does not communicate referential uncertainty accurately.[1]

What this means for people

  • A fluent AI response can conceal ambiguity, leaving people unable to recognise when the system has acted on the wrong object, record or instruction.
  • Interfaces should ask focused clarification questions and delay consequential actions when several plausible referents remain.
  • Warnings can also cause harm when poorly targeted, so products need evidence that uncertainty prompts improve decisions rather than merely adding cautionary language.

Global context

The human sample consisted of UK-based English speakers, while the authors are affiliated with Microsoft Research UK and Harvard University. The experiment does not establish outcomes in other languages, countries or professional settings; ambiguity may change with dialect, translation, specialist vocabulary and local interaction norms.

What the evidence does not yet show

  • This is a single unreviewed preprint with no independent replication and no second source validating the numerical results.
  • The 210-person study measured judgements about recorded puzzle turns, not completed actions in a live, consequential workflow.
  • Only three OpenAI model generations and one task design were tested, so the results do not estimate failure rates for other providers or deployed products.
  • The strongest hedge used oracle knowledge of whether the action was correct; the accessible manuscript provides no funding or competing-interest statement.

What to watch next

  • Independent replications using other model families, languages, task types and real collaborative decisions.
  • Whether non-oracle uncertainty signals can improve human intervention without generating excessive false alarms.
  • Product interfaces that expose alternative interpretations and require clarification before consequential action.

Evidence trail

Sources used for this report

Links checked 1 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Science & Research

Can AI design better optimisation rules?

A Nature Machine Intelligence paper reports that a structured LLM framework outperformed prior LLM approaches across 36 combinatorial-optimisation benchmarks and produced feasible algorithms for four new port-logistics problems. It remains an offline benchmark, not a live operations result.

8 min · 2 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.