Back to the news portal
Science & ResearchResearch paperResearchSource analysisSingaporeJapanUnited KingdomGlobal

Can AI design better optimisation rules?

A Nature Machine Intelligence paper reports that a structured LLM framework outperformed prior LLM approaches across 36 combinatorial-optimisation benchmarks and produced feasible algorithms for four new port-logistics problems. It remains an offline benchmark, not a live operations result.

By The Impact of AI Research DeskReleased 1 October 2026 at 22:56 BST8 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesLarge language modelsOperations researchAlgorithm designLogisticsReproducibility

Research topic

LLM-assisted design of complementary optimisation heuristics

The Impact of AI research cover asking whether AI can design better optimisation rules, with a problem contract branching into verified tools, heuristic portfolios, selection and evolution.
AI-generated editorial illustration. The contract and optimisation pathways are conceptual and do not depict measured results or a live routing, scheduling or packing system.

At a glance

  • 1LACE separates a verified problem contract from an LLM-driven search that builds and selects a portfolio of complementary heuristics.
  • 2Across 36 CO-Bench problems, the paper reports an average score of 0.945, versus 0.870 for the strongest compared LLM method and 0.571 for direct prompting.
  • 3Four new port-logistics problems tested structural generalisation, but all results remain offline and depend on formal specifications, runtime budgets and benchmark distributions.

Living evidence record

Impact record IAI-16YOTJ2

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

1 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

The research targets the slow craft of heuristic design

Routing vehicles, assigning crews, packing containers and scheduling machines are combinatorial problems: the number of possible solutions grows too quickly for exhaustive search. Practitioners therefore rely on heuristics that find good answers within an acceptable time without promising the mathematical optimum. Designing those rules normally requires domain knowledge, repeated testing and careful software engineering. The new paper asks whether a language model can automate more of that work without being trusted to improvise an entire solver in one step.

The authors' answer is LLM-driven Algorithm Construction via Complementary Evolution, or LACE. It does not simply ask a chatbot for code. It first defines a machine-checkable contract for the problem, then generates and evolves several heuristics whose strengths cover different instance types. That separation between a verified interface and a creative search process is the central contribution.[1][2]

A problem contract limits what generated code can misunderstand

LACE represents each task through an input schema, an output schema, a tool library and a heuristic portfolio. The schemas specify what the solver receives and must return. Verified tools check feasibility, calculate the objective and provide domain helpers. Generated heuristics can then focus on search strategy rather than rebuilding parsers and validators. Each contract module is smoke-tested before heuristics are admitted, according to the authors' implementation.

This architecture matters because a fluent model can produce syntactically plausible code that violates a constraint or returns an invalid data shape. A formal contract cannot guarantee a strong algorithm, but it can reject a large class of failures. It also makes human responsibility visible: someone still has to express the real objective, hard constraints, permissible data and evaluation rule. If that contract omits a safety or labour constraint, the optimiser will not restore it by intuition.[1][2]

The method searches for a team of specialists, not one champion

In its second stage, LACE uses seven operators to generate, combine and repair candidate heuristics under a fixed per-instance runtime budget. A mixed-integer selection step retains a portfolio of complementary specialists. One heuristic may work best on small or dense cases while another handles a different structure. The system scores the portfolio by whether its members collectively cover the distribution, instead of forcing a single average winner.

The released repository contains ten-heuristic portfolios for each task. Complementarity can improve benchmark coverage, but it adds a selection problem: at deployment time the system must know which heuristic to run or pay the cost of several. The benchmark's best-per-instance view should not be confused with an oracle available in live operations. Selection overhead, latency and failure handling are part of the real system even when the paper's algorithmic score is strong.[1][2]

The denominator is 7,109 instances across 40 problems

The main comparison uses 36 classical problems from CO-Bench spanning scheduling, routing, graph, packing, cutting, location, assignment and knapsack families. The authors then add four structurally new port-logistics problems covering port scheduling and several tugboat routing variants. The public manifest lists 7,109 test instances across all 40 problems, with the exact files and case indexes needed for reproduction.

That is a broad computational test bed, not a random sample of business decisions. Classical libraries are cleanly specified and have known objectives; real factories and ports contain missing data, changing priorities, human exceptions and constraints that may be negotiated rather than fixed. The four novel tasks are a useful guard against simple benchmark memorisation, but they were still created, formalised and evaluated by the research team.[1][2]

Scores improved substantially over compared LLM approaches

On the 36 classical problems, the paper reports a mean normalised score of 0.945 for LACE. The strongest existing LLM-based comparator scored 0.870, while direct prompting scored 0.571. The contrast is important because it attributes much of the gain to scaffolding—the contract, verified tools, evolution and complementary selection—rather than to language-model capability alone. Ablation studies reported by the authors also reduced performance when the tool library or portfolio mechanism was removed.

For the four new port-logistics problems, LACE reportedly scored between 0.97 and 0.99, while five LLM-based baselines failed to produce a feasible algorithm. Feasibility is a meaningful threshold: an elegant schedule that violates a berth or tugboat constraint is unusable. Yet a normalised benchmark score does not translate directly into money saved, delays prevented or emissions reduced. Those outcomes require comparison with established operations-research solvers and human-designed systems in live conditions.[1][2]

Open materials improve scrutiny, but generation still has a cost

The authors have released code, the 40 task contracts, instance data, evolved portfolios, per-instance results and reproduction scripts. The repository includes a no-API evaluation path for shipped portfolios, plus a Colab notebook. Re-running the generative search from scratch requires access to a language-model service. The default configuration names a commercial model endpoint, so reproduction of discovery costs money even though evaluation of the released outputs does not.

This distinction matters when assessing cost-effectiveness. An organisation can inspect and test the published heuristics cheaply, but designing new ones requires model calls, repeated execution and selection. The paper examines performance under runtime budgets; a procurement decision also needs total generation cost, engineering time, security review and maintenance when constraints change. Open code makes those questions answerable—it does not answer them automatically.[2]

Funding spans public agencies across three countries

The author team includes researchers linked to Nanyang Technological University and collaborators in Japan and the United Kingdom. The paper acknowledges support from Singapore's Agency for Science, Technology and Research and Ministry of Education, the Japan Science and Technology Agency, and the UK Engineering and Physical Sciences Research Council. The accessible publication metadata does not indicate a commercial sponsor for the central study, while the implementation relies on third-party language models that can change over time.

Model dependence is a reproducibility issue as well as a commercial one. Providers update weights, safety filters, prices and availability. The authors report robustness checks across different frontier-model backbones, but future users should record the exact model, prompt, seed where available, tool version and budget. Otherwise a nominal replication can be a different experiment.[1][2]

People are affected through the objectives humans choose

Better heuristics could reduce wasted journeys, idle equipment or scheduling time. They could also optimise an incomplete target that ignores fatigue, accessibility, job quality or resilience. A hospital schedule that minimises delay may create unsafe shift patterns; a logistics plan that cuts distance may concentrate noise in one community. The framework accelerates search over the encoded problem, which makes careful problem formulation more—not less—important.

Teams adopting generated heuristics should preserve human review of the contract, test hard constraints independently, compare against established baselines and monitor outcomes after deployment. Generated code belongs in the same assurance process as human code: version control, tests, security review, rollback and incident reporting. A high offline score is evidence for further evaluation, not permission to automate consequential decisions without oversight.[1][2]

What would change the assessment

The result would become more persuasive for industry if independent teams reproduced the full search, not only the released portfolios, and if blinded evaluations compared LACE with expert-built solvers on previously unseen operational data. Field trials should report feasibility violations, objective quality, runtime, generation cost, maintenance burden and the effect of changing constraints. Human experts should also assess whether the resulting heuristics are understandable enough to audit and repair.

Our assessment would weaken if gains disappear under fixed total compute, if the portfolio selector relies on hindsight, or if small contract errors create brittle behaviour. For now, this is strong peer-reviewed benchmark evidence that structure can make language models better algorithm-design assistants. It is not evidence that an autonomous system can safely optimise a port, factory, hospital or power grid without expert specification and operational validation.[1][2]

What this means for people

  • Faster heuristic design could improve routing, scheduling and packing where specialist optimisation expertise is scarce.
  • Workers and communities can be harmed when a formally valid objective omits fatigue, fairness, accessibility or resilience.
  • Open code allows technical teams to inspect the method, but accountable deployment still requires domain experts and affected people.

Global context

The research team and funders span Singapore, Japan and the United Kingdom, while the benchmark problems are drawn from internationally used operations-research collections. Applicability depends less on geography than on whether a local problem can be specified faithfully and tested against real constraints; the paper does not establish performance for any operating port, hospital or factory.

What the evidence does not yet show

  • The evidence comes from offline benchmark instances with formal objectives and constraints, not live operations.
  • Normalised scores do not directly measure financial savings, emissions, service quality or worker outcomes.
  • The four novel problems were designed and evaluated by the same research team.
  • Running the generative pipeline depends on changing third-party language models and incurs compute and API costs.
  • Portfolio performance and heuristic selection can add overhead that single-score summaries do not fully express.

What to watch next

  • Independent end-to-end replications using the released 7,109-instance manifest.
  • Head-to-head trials against expert operations-research solvers under equal total compute and time budgets.
  • Live pilots that disclose constraint violations, maintenance work and human overrides.
  • Whether generated algorithms remain robust when objectives, data distributions or provider models change.

Evidence trail

Sources used for this report

Links checked 1 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.