Can an AI choose your next experiment before you spend the compute?
A new preprint tests training intuition. We compare its narrow task with engineering, replication and architecture-search benchmarks to explain what research teams should measure next.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Whether predicting a promising training recipe translates into reproducible, cost-effective scientific work
At a glance
- 1The new study is an unreviewed benchmark, not a demonstration of autonomous discovery.
- 2Experiment choice, execution and replication require separate evaluation.
- 3Our proposed next test measures cost, failures and reproducibility on held-out projects.
Living evidence record
Impact record IAI-0KQRM3X
Evidence stage
Studied
Confidence
Supported
Reporting basis
Multi-source analysis
Independent support
Present
Record status
Monitoring
Last checked
1 October 2026
Source trail
4 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
What the new preprint actually measures
ArchitectureIQ, submitted by Zirui Ren and colleagues on 30 September, asks a language model to choose the best training recipe for a synthetic dataset before running it. Its static benchmark contains 500 questions, with outcomes established by execution. The authors report roughly 76% accuracy for frontier models against a random-choice baseline near 33%. Their human comparison involves ten researchers on a 50-question subset, so it should not be read as a universal machine-versus-scientist ranking. The paper identifies weaker performance on architecture-only questions and sensitivity problems around dataset properties. These are the authors' findings in an unreviewed preprint, rather than independently replicated results.[1]
Engineering a result is a different test
MLE-bench offers a useful earlier comparison because its tasks require agents to prepare data, train models and run experiments across 75 Kaggle competitions. It uses competition leaderboards to establish human reference points and examines resource scaling and possible training-data contamination. That design measures an extended engineering workflow rather than a single choice among recipes. Its original model scores date from 2024 and cannot rank today's systems.
Our interpretation is that a team should demand both measurements before handing over expensive experiments. A good recommendation can still fail during data preparation, package installation, checkpoint management or final evaluation. Conversely, an agent with modest initial intuition could become useful through careful execution and feedback. Procurement that collapses these stages into one impressive score would miss where staff still need to intervene.[2]
Replication asks whether the experiment can be trusted
PaperBench, released in 2025, asks agents to replicate 20 ICML 2024 papers from scratch. Its author-assisted rubrics divide the work into 8,316 gradable outcomes, and the evaluation includes an automated judge with a separate assessment of judging quality. The resulting replication score concerns implementation and empirical reproduction, rather than predicting a winner. It is another distinct capability test, with its own scope and potential grading errors.
For a research director, the practical question is whether a claimed result leaves enough evidence for someone else to challenge it. We would require the dataset version, split, random seeds, configuration, logs, failed runs and complete evaluation outputs. An attractive graph without that record should not unlock the next project stage. Junior researchers also need time to understand the reasoning, because checking an unfamiliar automated pipeline can become its own substantial task.[3]
Cheap benchmarking depends on what was precomputed
NAS-Bench-101 shows a different route to reducing research cost. Its 2019 dataset covers about 423,000 unique convolutional architectures, repeatedly trained on CIFAR-10, with more than five million trained-model records. Researchers can query those stored results instead of paying for every trial again. The study made comparisons easier to reproduce within a deliberately bounded search space; it did not establish a universal recipe for every task.
The comparison suggests a practical division of labour. Use stored experimental evidence where it matches the problem, ask an assistant to propose alternatives where it does not, and run a small controlled pilot before scaling. A benchmark's convenience becomes a weakness if its task is far removed from a laboratory's data, hardware or objective. A company predicting medical outcomes, for example, should not infer clinical readiness from success on a synthetic training puzzle.[4]
A proposed pilot for real research teams
The newsroom's proposed next test is a prospective comparison on projects selected before any model sees their results. Give an assisted team and an unaided team the same data access, hardware budget and stopping rules. Record the best independently verified result, total compute, researcher hours, abandoned runs and time spent repairing or checking output. Include projects where a simple established method is already strong, so extra complexity must earn its cost. This is an evaluation proposal, not an outcome established by the cited studies.
Teams should also record which recommendations were rejected and why. If a system repeatedly proposes experiments that violate a resource limit, rely on unavailable data or optimise the wrong objective, counting only its successful suggestions would give a misleading picture. Publish results for the complete project denominator and separate assistance from end-to-end autonomy. Repeating the trial in another institution would help show whether the advantage comes from the system or unusually experienced operators.
This matters for access as well as productivity. A smaller university may value avoiding a few wasted runs more than achieving a marginally better leaderboard score. A well-funded laboratory may instead care about repeatable coordination across many projects. Those are different buying decisions and need different evidence. For now, the strongest practical response is to test AI as a documented experiment adviser, then expand its responsibility only where measured work supports that decision. A pilot should document the information boundary too. Give each participant an explicit list of permitted files and tools, record external requests, and separate confidential project data from public benchmark material. Review whether a proposed experiment depends on information that the team is entitled to use. Otherwise an apparent productivity gain could simply reflect unequal access to data, hidden human assistance or a workflow that another institution cannot reproduce under its own constraints.[1][2][3][4]
What this means for people
- Researchers could test whether documented AI advice saves scarce compute and supervisory time.
- Students and junior staff should retain access to the experiment record and the reasons for decisions.
Global context
A useful cross-border comparison would hold task access and resource budgets constant while recording differences in infrastructure and researcher support. Benchmarks cannot establish those institutional effects by themselves.
What the evidence does not yet show
- ArchitectureIQ is a preprint; independent replication has not been established here.
- The comparison benchmarks concern different tasks, budgets and historical models; their percentages are not interchangeable.
- The proposed pilot has not been run by this newsroom.
What to watch next
- Independent replication and matched human comparisons on new tasks.
- Prospective evidence of lower total research cost without weaker reproducibility.
Evidence trail
Sources used for this report
Links checked 1 October 2026
This report is labelled multi-source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
AI agents often abandoned the plans they declared; stronger execution structure helped in benchmarks
New analysis today of a 29 September preprint led from Abu Dhabi: generic agents often failed to preserve their stated reasoning pattern. Structured executors improved some benchmark results, but the study is unreviewed and its fidelity labels received only limited human validation.
6 min · 2 sources
Science & Research
Can AI plan under pressure? New Game Arena research tests chess, poker and Werewolf
A newly listed preprint describes an open evaluation arena where models face other models in games with different kinds of uncertainty. It tests strategic behaviour, not whether an agent is safe to run a business or make decisions for people.
4 min · 1 source
Science & Research
AI research agents generate thousands of ideas—but may converge on the same directions
A preprint comparing research-agent frameworks generated 37,802 ideas from shared literature and found evidence that automated systems can narrow exploration even while increasing output.
4 min · 1 source
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.