Back to the news portal
AI Risks & SafetyResearch paperResearchSource analysisLatin AmericaEuropeInternational

Can hashes make AI conversations auditable without publishing them?

A peer-reviewed experiment converted nearly four million public chatbot interactions into cryptographic commitments and detected eight induced ledger manipulations. It is a promising integrity mechanism, not proof that a conversation is true, authorised or safe.

By The Impact of AI Editorial DeskReleased 1 October 2026 at 19:24 BST9 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesAI auditPrivacyCryptographyGovernanceLarge language models

Research topic

Whether public LLM conversations can be transformed into privacy-aware cryptographic records that support scalable integrity checks without placing the conversation text in a ledger

At a glance

  • 1The researchers transformed 1,870,989 source records from LMSYS-Chat-1M, WildChat and Chatbot Arena into 3,971,887 auditable prompt-response or comparison units.
  • 2A controlled 139,258-record ledger detected all eight induced manipulations across five predefined categories, while an unmodified baseline produced no error signals.
  • 3The method can show that committed records changed, but only relative to a trusted checkpoint; it cannot prove the original text was truthful, authorised or produced by the claimed model.

Living evidence record

Impact record IAI-0S2VM2W

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

1 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Frontiers in Artificial Intelligence published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

The research question is narrower than 'can blockchain make AI trustworthy?'

The study asks whether irregular chatbot conversations can be turned into stable evidence that an auditor can later verify without copying the prompts and answers into a public ledger. The authors begin from a practical problem: a chat log may contain several turns, alternative model responses, votes, changing metadata and sensitive text. A hash is only reproducible if everyone hashes the same representation. The first task is therefore canonicalisation—deciding exactly what counts as the record and serialising it deterministically—before any cryptographic claim can be meaningful.

Researchers from universities in Ecuador, Chile and Spain tested a pipeline on three public sources: LMSYS-Chat-1M, WildChat and Chatbot Arena. LMSYS contributed 1,000,000 source conversations, WildChat 837,989 records and Chatbot Arena 33,000 comparisons. After filtering 43,148 structurally incomplete records, the pipeline produced 3,971,887 auditable units: 1,989,603 from LMSYS, 1,943,026 from WildChat and 39,258 from the arena comparisons. The denominator is units after transformation, not unique people, models or real production decisions.[1]

How the proposed audit trail works

Each canonical unit retains identifiers and traceability fields while replacing prompt, response, metadata and decision content with SHA-256 or keyed HMAC-SHA-256 commitments. The records are ordered, divided into blocks and combined through Merkle roots. Each block includes the previous block's hash, forming a lightweight chain, while a separate membership map allows an auditor to test whether a particular record belongs to the committed set. The original conversational text stays off-chain; the ledger represents commitments to it rather than the content itself.

That separation matters for privacy, but it is not anonymity by magic. A plain hash of predictable text can sometimes be guessed by hashing possible inputs. The HMAC mode adds a secret key, making such guesses much harder for someone without the key, but it also creates a governance dependency: an institution must protect, rotate and revoke keys and decide which auditors may use them. The paper treats those production controls as outside the experiment rather than claiming to have solved them.

The algorithm also makes an unusually important distinction. If an auditor has an independently preserved or externally anchored checkpoint, rewriting the later ledger can be detected against it. Without that checkpoint, the system can only show that the supplied ledger is internally consistent. An attacker able to rewrite every stored artifact could create a new consistent chain. The framework therefore provides tamper evidence relative to trusted evidence; it does not create an independent history from nothing.[1]

What the experiments measured

The performance tests used loads of 10,000 to 500,000 normalised interaction records, each combined with the fixed set of 39,258 Chatbot Arena comparison units. Total audit records therefore ranged from 49,258 to 539,258. Mean end-to-end pipeline time rose from 7.393 seconds at the smallest load to 95.090 seconds at the largest, averaged over five runs. Full record-set verification stayed around 132,000 to 142,000 records per second through the 289,258-record scenario, then fell to roughly 102,000 records per second at 539,258 records. The authors describe the overall timing as near-linear over the tested range.

On a fixed 139,258-record pilot set, HMAC-SHA-256 reduced audit-record generation throughput from about 6,805 to 5,843 records per second compared with unkeyed SHA-256, a 14.14% penalty. Verification throughput was effectively unchanged within run-to-run variation. Five runs in each mode completed successfully. Those are engineering measurements on the authors' implementation, not service-level guarantees for a bank, hospital or government system with authentication, networking, redundancy and concurrent users.

The storage accounting is also specific. An 11,899.55 MiB normalised text repository became 1,473.44 MiB of hash-only audit records. Block headers alone occupied 2.03 MiB; the operational ledger plus membership structures needed for record-level verification occupied 380.52 MiB. This is evidence that commitments and indexes take less space than the source text. It should not be described as lossless compression: reconstructing a conversation still requires the separately preserved off-chain content.[1]

The tamper test succeeded inside a controlled threat model

The authors introduced eight manipulations across five categories: altered decision or Merkle hashes, broken hash-chain links, reordered blocks, missing membership records and inconsistent dataset membership. The complete reconstruction procedure detected all eight, with no false acceptance. Three clean-ledger validations produced zero error signals. Median validation times for the manipulated categories were between roughly 609 and 627 milliseconds on the fixed 139,258-record, 140-block ledger.

This is a functional test, not a population estimate of attack detection. There were only eight deliberately designed manipulations, so the result does not supply a confidence interval for every adversarial technique. It also does not test stolen HMAC keys, compromised logging software, false records committed at ingestion, selective non-logging, forged model identities or an administrator rewriting both the ledger and every checkpoint. Those failure modes sit before or around the cryptography rather than inside the tested chain.

The study similarly does not evaluate whether any chatbot answer was correct, fair, safe or lawful. Cryptographic integrity can preserve a bad decision just as faithfully as a good one. Auditors would still need policies defining which events must be recorded, authenticated metadata identifying the system and operator, access controls around the off-chain text, retention limits, and substantive tests of outcomes. Integrity is one layer of accountability, not accountability itself.[1]

Practical implications for people and institutions

For a person challenging an automated insurance decision, clinical recommendation or public-service interaction, a reliable record could establish which prompt, response, model identifier and policy state were committed at the time. That can reduce disputes about whether a log was quietly edited after harm occurred. Keeping the text off-chain may also reduce unnecessary exposure compared with publishing complete interactions. But access to the original record remains essential: a hash alone cannot explain what the system said or allow a person to contest it.

For organisations, the paper suggests a design discipline rather than a product ready for purchase. Teams must define the canonical event before deployment, link it to authenticated software and configuration records, anchor checkpoints beyond the control of a single administrator, and test recovery when keys or off-chain data are lost. They also need data-protection analysis because public research datasets can contain user-submitted personal information; the authors say they did not reproduce prompts or try to re-identify users, but real deployments would face stronger consent and purpose-limitation obligations.

The international author team and public datasets make the mechanism broadly inspectable, yet law and institutional capacity vary. A regulator may accept an internal signed checkpoint in one jurisdiction and require external timestamping or certified records in another. Organisations in lower-resource settings may benefit from the relatively light ledger, while facing harder questions about secure key custody and durable storage. Standards work should therefore test interoperability and redress, not just hashing speed.[1]

Funding, interests and what would change the assessment

The authors report no financial support and no commercial or financial conflict of interest. They disclose using generative AI for language refinement and editorial improvement, not for data generation, experiments, analysis or conclusions. The data sources are public and linked, although the article's data statement directs further requests to the authors rather than identifying a public repository for the complete implementation and experimental artifacts.

Confidence would rise with independent reproduction using released code, frozen data versions and hardware details; adversarial tests designed by outside security teams; and deployment studies in which a separate institution controls the checkpoint. Evaluations should add compromised loggers, stolen keys, missing-event attacks and attempts to correlate hashes with sensitive text. A legal or operational pilot should also measure whether the record helps affected people obtain an explanation or correction, not merely whether an engineer can validate a Merkle root.

The evidence supports a precise conclusion. The researchers demonstrated that millions of heterogeneous public chat interactions can be represented as compact commitments, and that a prototype can detect specified changes efficiently relative to preserved evidence. They did not demonstrate an immutable production blockchain, prove upstream provenance or show better decisions. The next assessment should turn on independent replication and whether the mechanism survives hostile governance conditions while improving real audit and redress outcomes.[1]

What this means for people

  • A tamper-evident log could help people prove which AI interaction informed a consequential decision, provided the original record remains available for challenge.
  • Keeping conversation text off the ledger reduces direct exposure but does not remove the need for secure storage, access control and deletion rules.
  • Auditors and frontline workers would need tools that translate cryptographic checks into understandable evidence rather than adding another opaque compliance layer.

Global context

The authors are based in Ecuador, Chile and Spain, and the datasets aggregate public conversations from internationally used systems. That gives the work wider geographical roots than many AI-audit studies, but the experiment does not report performance by language, region or legal regime. Deployment would need locally valid checkpointing, data-protection, labour and redress arrangements rather than assuming one technical ledger satisfies every jurisdiction.

What the evidence does not yet show

  • The eight manipulation scenarios are controlled functional checks, not a comprehensive adversarial security evaluation.
  • Without a trusted external checkpoint, the method demonstrates internal consistency rather than tamper evidence against an attacker who can rewrite all artifacts.
  • Experiments use public research datasets and a lightweight simulated ledger, not a production system with institutional identity, access and key-management controls.
  • Cryptographic integrity does not establish factual accuracy, fairness, legal authority, completeness of logging or the authenticity of upstream model provenance.

What to watch next

  • Independent replication with released code, pinned datasets and reproducible hardware and timing details.
  • Red-team tests covering stolen keys, compromised loggers, omitted events and full-ledger rewriting.
  • External checkpointing trials in regulated services and evidence that records improve explanation, challenge and correction for affected people.
  • Privacy analysis of linkability, retention and access to off-chain conversations across different legal regimes.

Evidence trail

Sources used for this report

Links checked 1 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

AI Risks & Safety

Does the White House AI accord create enforceable safety rules?

Six major AI companies signed a four-layer commitment covering internal controls, external evaluation and board oversight. The text is concrete enough to audit later—but voluntary, undefined and silent on publication, deadlines and sanctions.

5 min · 4 sources

AI Risks & Safety

OpenAI holds GPT-6.1 Astra release after safety tests fall short

The company confirmed on 28 September that the planned October launch would not go ahead. Reuters and AP report concerns about scope, authorization and how the model describes its actions; detailed test results remain private.

4 min · 2 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.