AI agents often abandoned the plans they declared; stronger execution structure helped in benchmarks
New analysis today of a 29 September preprint led from Abu Dhabi: generic agents often failed to preserve their stated reasoning pattern. Structured executors improved some benchmark results, but the study is unreviewed and its fidelity labels received only limited human validation.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
When agents are allowed to revise a plan, which execution constraints improve successful adaptation while preserving a verifiable record of user authorisation and material deviations?
At a glance
- 1The primary evidence was submitted on 29 September and is an unreviewed preprint; this article is new analysis today, not a claim of a 30 September source.
- 2Generic Plan-and-ReAct agents preserved their declared structure in about 22% to 45% of reported runs, while pattern-specific executors improved results on some controlled benchmarks.
- 3A readable plan is not an audit trail: people need action logs, deviation notices and renewed permission before consequential changes.
Living evidence record
Impact record IAI-1QTYVOU
Evidence stage
Studied
Confidence
Supported
Reporting basis
Multi-source analysis
Independent support
Present
Record status
Monitoring
Last checked
30 September 2026
Source trail
2 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
New analysis today — not a new source today
A preprint led by ADIA Lab in Abu Dhabi asks a deceptively simple question: when a large-language-model agent announces how it will approach a task, does its subsequent behaviour preserve that plan? The paper was submitted to arXiv at 17:47 UTC on 29 September and appeared in arXiv's 30 September announcement list. This article is new analysis today; the primary evidence is dated 29 September. The manuscript has not been peer reviewed, independently replicated or tested in a production service.
The authors compare generic planning with executors designed around specific reasoning patterns. Their central claim is not that agents never plan. It is that a declared pattern can dissolve once execution begins: a system may say it will work step by step, branch through alternatives or search a tree, yet later produce actions that no longer match that structure. That distinction matters because a readable plan can give users and reviewers a false sense of control if the system's actual tool use follows a different process.[1]
What the study measured
The evaluation covers four controlled benchmarks: 1,341 Mind2Web tasks, 204 WebArena tasks, 500 SWE-bench Verified software issues and 134 ALFWorld household-simulation tasks. The researchers tested Qwen3.6-35B-A3B, DeepSeek-V4-Flash and Gemma-4-26B-A4B-it, with repeated evaluations using three random seeds. The tasks span web navigation, software repair and simulated embodied action. They do not form a representative sample of banking, healthcare, public services or everyday consumer use.
For each planning pattern, the team specifies an execution interface intended to preserve its structure. A plan-and-execute agent creates a plan and follows its steps; a tree-search executor keeps branches and scores alternatives; a debate-style system separates proposing, criticising and deciding. The researchers then use task-specific success checks and an automated verifier to assess whether trajectories retained the declared pattern. This measures structural fidelity separately from whether the final answer or action was correct.[1]
What changed when execution was structured
Across three benchmarks where the comparison was reported, a generic Plan-and-ReAct approach preserved the declared structure in only about 22% to 45% of runs, with lower fidelity on longer plans. Pattern-specific executors improved task success in some settings: on ALFWorld the reported rate rose from 0.48 to 0.92, and on SWE-bench Verified from 0.36 to 0.44, relative to Plan-and-ReAct. These are benchmark outcomes under the authors' harness, not estimates of the reliability of deployed assistants.
More computation is a potential confound because search can try several candidate paths. The authors include a retry-matched control to compare structured search with spending a similar budget on repeated attempts. Search retained an advantage on ALFWorld and WebArena, but the benefit was not uniform across tasks and models. The study also found that an agent's declaration did not reliably select the best reasoning pattern task by task: a task-blind permutation test encompassed the observed differences, with the smallest reported one-sided p-value at 0.094.[1]
Impact on people and global context
For people, the immediate lesson is about interface honesty. A displayed plan should not be treated as an audit trail unless the product records how each later action relates to it. In workplace software, an agent could present an orderly checklist while its tool calls skip approval steps or follow another route. Developers can reduce that gap by enforcing execution constraints, logging deviations and requiring renewed permission when a plan changes, especially before actions affecting money, records, employment or access to services.
NIST's May review of public responses about AI agents describes novel security risks as a barrier to adoption and says existing cybersecurity practice may need adaptation. That policy context supports closer attention to agents that act through tools, but it does not validate the new paper's numerical results. The research is geographically notable because the lead institution is in the United Arab Emirates, with collaborators in Spain, Luxembourg and the United States; the tested models and benchmarks still do not establish country-specific outcomes.[1][2]
Evidence, limitations and what to watch
The fidelity verifier received a limited human check on 100 ALFWorld trajectories. Reported agreement was only fair: Cohen's kappa was 0.35 between human raters and 0.31 and 0.34 between the verifier and each rater. That uncertainty matters because the paper's structural-fidelity claims depend on classification. The accessible manuscript does not provide a funding or conflict-of-interest statement, so readers cannot assess those disclosures from this version. The authors say code and trajectories will be released upon publication; independent reviewers should verify availability before treating the work as reproducible.
Structural fidelity is not equivalent to safety or factual accuracy. An agent can follow a plan faithfully and still have a bad objective, use wrong information or cause harm. Conversely, changing a plan can be sensible when new evidence appears. The next useful studies should test whether enforced structures improve calibrated success without blocking justified adaptation, using external evaluators, larger human-labelled samples and realistic high-consequence workflows. Until then, the paper supports a narrower conclusion: declarations alone are weak evidence that an agent will execute as described.[1][2]
What this means for people
- Users should be shown when an agent departs from its declared plan, especially before consequential tool actions involving money, personal records, employment or access to services.
- Organisations need logs that connect each action to the approved plan and distinguish justified adaptation from unauthorised scope expansion.
Global context
The lead institution is in Abu Dhabi and the author group spans the UAE, Spain, Luxembourg and the United States. The benchmarks and three tested model families are international in origin, but they do not measure population-level or country-specific effects. NIST's agent-security review offers policy context rather than independent confirmation.
What the evidence does not yet show
- This is an unreviewed preprint tested on controlled benchmarks, not on deployed services or real people; no independent replication is available.
- Human validation covered only 100 ALFWorld trajectories and produced fair agreement, so automated structural-fidelity labels carry material uncertainty.
- Structured search can use more computation, reported gains were not uniform, and following a pattern does not prove factual accuracy, safety or user benefit.
- The accessible manuscript contains no funding or conflict-of-interest statement, and promised code and trajectories should be checked when released.
What to watch next
- Independent replication with released code, external evaluators and larger human-labelled samples across realistic tool-using workflows.
- Product evidence showing whether enforced execution structures reduce unauthorised actions without preventing sensible plan changes.
- Interfaces that disclose deviations and request renewed consent before a changed plan causes a consequential action.
Evidence trail
Sources used for this report
Links checked 30 September 2026
This report is labelled multi-source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
People accepted 78% of wrong AI actions when uncertainty stayed hidden
New analysis today of a 30 September preprint: in a controlled puzzle study, people often approved incorrect AI moves when the system did not reveal ambiguity. Targeted warnings helped, but the best-performing warning relied on oracle knowledge that a real product would not have.
6 min · 1 source
Science & Research
Can an AI choose your next experiment before you spend the compute?
A new preprint tests training intuition. We compare its narrow task with engineering, replication and architecture-search benchmarks to explain what research teams should measure next.
5 min · 4 sources
Science & Research
Microrobot navigation can be trained in minutes in new study; patient use remains untested
A peer-reviewed Hong Kong-led paper reports under-ten-minute policy training across thousands of simulated vessel environments and controlled robot tests. It does not show a clinical procedure or patient benefit.
4 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.