Back to the news portal
TechnologyPrimary sourceAnalysisSource analysisUnited StatesInternational

Does GPT-6.1 Sol make AI work cheaper? Count successful tasks, not just tokens

OpenAI's 29 September release reports stronger performance at lower task costs in selected evaluations. The useful comparison for buyers is the cost of checked, completed work in their own setting.

By The Impact of AI Editorial DeskReleased 30 September 2026 at 07:09 BST4 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesModel evaluationCostReliabilityAI procurement

At a glance

  • 1Selected vendor tests suggest improved capability; local total-cost testing is still needed.
  • 2Safety capability classifications are not estimates of harm probability.

Living evidence record

Impact record IAI-1NB6TYB

Explore the full tracker

Evidence stage

Announced

Confidence

Supported

Reporting basis

Source analysis

Independent support

Not yet

Record status

Monitoring

Last checked

30 September 2026

Source trail

2 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

What OpenAI released

OpenAI announced GPT-6.1 Sol on 29 September 2026 for ChatGPT Work, Codex and its API. The launch reports improvements on software, professional-document and computer-use evaluations. Standard API pricing is listed at $2 per million input tokens, $0.10 for cached input and $10 for output. Those prices describe consumption, not the complete cost of delivering a business result.

One concrete comparison concerns factual errors on deliberately difficult conversations selected because users had flagged an earlier model's mistake. At low reasoning effort, OpenAI reports a reduction from 11.4% for GPT-6 Sol to 7.7% for GPT-6.1 Sol. The denominator is evaluated answers in that selected test, not all everyday user queries. The launch does not supply the sample size alongside that headline comparison. These results are vendor-reported and should not be presented as an independently measured error rate for ordinary work.[1]

What the safety documentation adds

The accompanying system-card addendum is also dated 29 September. OpenAI classifies the model as Critical for cybersecurity capability and High for biological and chemical capability under its own framework, applying the GPT-6 Astra safeguards stack. Those labels concern assessed capability; they are not probabilities of a harmful event. The document warns that research or API evaluations may differ from production behaviour because prompts, tools and reasoning settings differ.

Our assessment is that performance and permissions should be evaluated together. A system that finishes more tasks can still create avoidable costs if its output requires correction, if it acts beyond scope or if it does not disclose a failed information search. A business choosing a cheaper model should therefore retain the checks appropriate to the action. Price does not establish trustworthiness, and a model upgrade does not automatically validate the workflow around it.[2]

How to compare the cost of useful work

For a document-processing pilot, a team could select a representative batch before seeing either model's answers. Include awkward tables, missing attachments and questions that cannot be answered from the documents. Define which mistakes would make a result unusable. Then record accepted outputs, human review minutes, retries and escalations. The relevant denominator is the number of correctly completed cases. Counting all attempted cases as productivity would reward a system for generating work that somebody must redo.

The comparison should also keep task difficulty and tool access consistent. If one configuration can use a search service or a larger context window while another cannot, the experiment is comparing systems rather than only models. That can be a legitimate purchasing question, but it should be described honestly. Record the configuration and date, because changing a prompt or enabling an additional tool may alter the result more than a headline model ranking suggests.

Why the conclusion remains local

For staff, a lower operating cost could make assistance available for more routine tasks. It could also encourage organisations to process more work without increasing the capacity to review exceptions. Our recommendation is to set a review budget before expanding volume, and to preserve a route for workers to challenge an apparently confident answer. Those organisational choices determine whether cheaper inference reduces effort or creates a larger checking queue.

The launch is relevant internationally, but document languages, local terminology and professional rules vary. A result on a shared benchmark cannot establish equivalent performance on every country's records. Our assessment would change with independent, reproducible comparisons that include total cost and consequential errors across those settings. Buyers can already test a bounded workflow, but should report their own sample and acceptance criteria rather than borrowing a vendor's percentage as a guarantee.

What this means for people

  • Staff time spent correcting results belongs in any claimed productivity saving.

Global context

Country, language and workflow differences require local evaluation.

What the evidence does not yet show

  • The launch's selected factuality comparison does not report its sample size beside the headline.
  • Research settings and production systems can behave differently.

What to watch next

  • Independent comparisons including retries, human checking and completed-case accuracy.

Evidence trail

Sources used for this report

Links checked 30 September 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Technology

What does beating a Stratego champion prove about AI under hidden information?

Ataraxos won 15 of 20 games against one of Stratego's most decorated players and transferred its methods to three other games. The peer-reviewed Nature paper shows a real advance in efficient game strategy—not that the system can already run negotiations, markets or military decisions.

6 min · 3 sources

Technology

Does self-hosting keep an AI coding agent away from sensitive source code?

IBM has made a customer-managed version of its Bob development agent generally available, including supported air-gapped deployments. Local inference can keep code inside an approved environment, but the announcement does not prove that every integration, model or generated change is secure.

6 min · 2 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.