Back to the news portal
Society & MediaResearch paperResearchSource analysisUnited StatesNorth America

Will patients avoid doctors who say they use AI?

A preregistered experiment with 1,030 US adults found lower ratings and appointment intentions for fictional family doctors who said they used AI. The effect is relevant to patient trust, but an advert-based intention is not a real healthcare choice.

By The Impact of AI Editorial DeskReleased 1 October 2026 at 22:01 BST10 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesPatient trustPrimary careAI disclosureHealth communicationConsumer behaviour

Research topic

Whether disclosing AI use changes how US adults evaluate fictional family doctors and whether those reactions vary with physician or participant demographics

The Impact of AI research cover asking whether patients will avoid doctors who say they use AI, with conceptual doctor-profile cards and a choice between booking and seeking a second opinion.
AI-generated editorial illustration. It represents a fictional choice experiment and does not depict a doctor, patient or study participant.

At a glance

  • 1The preregistered study analysed 1,030 US adults who rated fictional family-doctor advertisements stating either that the doctor uses AI or never uses AI for diagnostics and therapy.
  • 2AI-use disclosure reduced ratings of warmth and competence and lowered several stated behavioural intentions, but the outcomes were hypothetical responses rather than observed appointments or treatment adherence.
  • 3Physician age, gender and race did not significantly change the main penalty after multiple-comparison correction; participant-age and race patterns were exploratory and require confirmation.

Living evidence record

Impact record IAI-1VTTMAD

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

1 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

The finding concerns a disclosure cue, not the quality of care

Researchers from the University of Würzburg, the University of Cambridge, the University of Queensland and Charité asked whether saying that a family doctor uses artificial intelligence changes how that doctor is perceived. Their peer-reviewed brief communication, published on 1 October, reports a preregistered online experiment in which 1,030 US adults evaluated fictional medical-practice advertisements. Doctors presented as AI users received lower ratings and weaker stated willingness to book than doctors presented as never using AI.

The experiment does not compare care delivered with and without an AI tool. Participants were not shown a diagnosis, error rate, clinical outcome, explanation or product. The wording said that the fictional doctor used AI tools for diagnostics and therapy; the control said the doctor never used them. The result therefore measures reaction to a broad disclosure cue under limited information. It cannot establish whether a real patient would reject a useful tool, whether the tool improves care or whether distrust is justified in a particular case.

That distinction is central to the practical importance of the study. Health systems increasingly debate when patients should be told that AI supported documentation, triage, imaging or treatment decisions. Disclosure can protect autonomy, yet the phrase 'uses AI' may cover tasks with very different risks and degrees of clinician control. The paper shows that the wording itself can influence social judgement. It does not support hiding AI use or treating scepticism as irrational.[1][2]

How 1,200 recruits became the 1,030-person analysis

Data were collected through Prolific in November 2025. The researchers used the platform's representative-sample feature to approximate the US adult population across age, binary gender and registered race categories. An advance calculation suggested that 800 people would be needed for the planned comparison, and 1,200 were recruited to allow for drop-out and uneven groups. The final analysis excluded 170 people—14%—who failed at least one of three manipulation checks at the end of the survey.

The resulting sample included 520 women, 487 men, 18 non-binary participants and five people who did not state a gender; mean age was 46.9 years. Participants reported 674 White, 123 Black, 96 Hispanic, 72 Asian and smaller other categories. More than half had a university degree. Repeating the main analyses with the excluded participants retained led to virtually identical conclusions, but exclusion after multiple end-of-study checks remains important context for interpreting the denominator.

A stratified online panel is not identical to a probability sample or a waiting-room population. Prolific participants may be more experienced with online research and digitally literate than the wider public, and the authors say attitudes to AI in this sample were relatively positive. The result applies most directly to US adults responding to an online scenario in late 2025; it should not be converted into a global estimate of patient distrust.[1][2]

The experiment varied AI wording and three physician characteristics

The design crossed physician gender, age and AI-use statement between participants and physician race within participants. Each person was randomly assigned to one of eight combinations—female or male, relatively younger or older, and uses AI or never uses AI—and then reviewed three adverts in randomized order for fictional White, Black and Asian doctors. Participants therefore saw only one AI-use condition but compared three race presentations within it.

The adverts used realistic physician images generated with Google's Gemini 2.5 Flash Image. Prompts and vignettes are available in the study repository. The authors note two possible confounds: participants may have recognised the images as synthetic, and minor visual details such as lighting or background people were not identical. Racially distinctive fictional names were also used to reinforce the intended categories, potentially introducing cultural signals beyond the planned manipulation.

After viewing each advert for 20 seconds, participants rated seven warmth items and six competence items, both producing high internal-consistency scores. They then rated willingness to choose the doctor, recommend the doctor, follow a treatment plan, book again or seek a second opinion. Numeric questions asked how far they would travel, how much they would pay and how long they would wait. These are attitudinal and stated-intention outcomes; no appointment was offered and no money or treatment was at stake.[1][2]

The main penalty was consistent, while its size varied

In the preregistered analyses, fictional doctors who said they used AI were rated as less warm and less competent. The reported partial eta-squared values were 0.04 for warmth and 0.05 for competence. Appointment willingness, treatment-plan adherence and willingness to book again each showed an AI-use effect of 0.07; willingness to recommend showed 0.06. Participants were also more inclined to seek a second opinion, with an effect of 0.03.

The numeric trade-offs moved in the same direction: respondents said they would travel less far, pay less and wait less long for the AI-using doctor. Those effects were smaller—partial eta-squared values of 0.03, 0.01 and 0.02 respectively—and the researchers applied non-preregistered outlier rules to these numeric measures after extreme answers, including a reported travel time above 10,000 minutes. Supplementary results without those removals are available.

Statistical significance is not the same as a large practical effect, especially in a sample exceeding one thousand. The experiment consistently detected a negative cue across several related outcomes, but it did not show the proportion of people who would definitely refuse care or the number of real appointments lost. The outcomes are also correlated: a person who sees the same fictional doctor as less competent may predictably lower several intentions. They should not be counted as separate independent replications.[1]

Demographic findings need two different confidence labels

The preregistered tests found no significant interaction between AI use and the fictional physician's age, gender or race after Benjamini-Hochberg correction for multiple comparisons. Within this design, the broad AI-use cue outweighed those manipulated doctor characteristics. That is a useful negative result: it gives no strong evidence that one displayed physician group uniquely bears the penalty.

Participant demographics were examined later in exploratory analyses that were not preregistered and did not receive the same multiplicity correction. Adults under 45 showed a stronger AI-use penalty on most measures, but the study did not directly test why. Familiarity with AI could increase awareness of limitations, as the authors suggest, or the split could capture other age-related differences. It should be treated as a hypothesis for confirmation rather than a targeting rule for healthcare communication.

A small three-way pattern involving physician and participant race appeared for warmth, with partial eta-squared of 0.01 and an unadjusted p value of 0.044. The researchers explicitly caution that the finding may reflect the within-person race manipulation and did not measure the underlying mechanism. It would be inappropriate to design race-specific messages from this single exploratory result. Replication with independently validated images and pre-specified comparisons is needed first.[1]

What this means for patients, doctors and health services

Patients need more than a generic AI label. A useful disclosure should say what the system does, what data it uses, whether a clinician reviews the output, what important limitations are known and who remains responsible for the decision. A tool that formats notes is different from one that recommends treatment. Grouping both under 'uses AI for diagnostics and therapy' may create an information vacuum that people fill with their own expectations.

Doctors and service leaders should not respond by minimising or concealing meaningful AI involvement. The experiment did not compare alternative explanations, so it does not show that transparency inevitably reduces trust. A future study could test whether a concrete account of clinician oversight, patient choice and verified performance changes the reaction. Communication should remain accurate: reassurance unsupported by local safety evidence would replace one trust problem with another.

For healthcare organisations, the relevant operational measures are actual appointment choices, questions raised during consent, second-opinion requests, complaints, opt-outs and whether disclosure affects different groups unequally. Those outcomes should be assessed alongside accuracy, time saved and staff workload. Trust is not a marketing obstacle to be optimised away; it is part of whether patients can understand and challenge the systems involved in their care.[1]

Funding, openness and what would change the assessment

The study was supported by the Faculty of Humanities at the University of Würzburg, with open-access funding through Projekt DEAL. One author reports employment with Pfizer outside the submitted work; the other authors declare no competing interest. The hypotheses, analysis plan and target sample were preregistered. Vignettes, underlying data and the analysis script are linked through the Open Science Framework, which materially improves scrutiny and reproducibility.

Confidence would rise with a preregistered field experiment in real primary-care services where patients encounter a specific, accurately described AI use and make an observable choice. Useful comparisons would include no disclosure, a bare disclosure, a detailed explanation of purpose and oversight, and a non-AI decision-support tool. Researchers should measure actual booking, continued care, comprehension and clinical safety across languages and health systems.

The assessment would weaken if the penalty disappears when people know the doctor's record, understand the tool or make a consequential choice rather than rate an advert. It would strengthen if similar effects recur in real clinics and remain after verified performance and clinician accountability are explained. For now, the study provides good evidence that a salient, broad AI-use statement can lower US consumers' evaluations of a fictional doctor—not that patients generally reject AI-supported medicine.[1][2]

What this means for people

  • Patients need understandable explanations of what an AI system does and who remains responsible for their care.
  • Doctors may face a trust penalty from a vague AI label even when the tool's role, benefit and risk differ substantially.
  • Health services should measure real choices and comprehension rather than treating favourable survey ratings as acceptance.

Global context

The experiment used a US online sample and US-style family-doctor advertisements. It cannot quantify attitudes in the UK, Europe, Asia, Africa or the Middle East, where primary-care access, payment, waiting times and trust in institutions differ. The transparent methods make replication feasible, but each health system needs local evidence and locally meaningful disclosure wording.

What the evidence does not yet show

  • The outcomes are ratings and hypothetical intentions after fictional adverts, not observed appointments, treatment adherence or clinical outcomes.
  • The AI-use wording was broad and salient, contrasting 'uses AI' with 'never uses AI' without identifying a task, product, performance level or oversight process.
  • The Prolific sample approximated selected US census categories but was an online panel with relatively high education and positive AI attitudes.
  • Fourteen per cent of recruits were excluded for failing at least one manipulation check, although the full-sample sensitivity analysis gave similar conclusions.
  • Participant-age and participant-race patterns were exploratory, not preregistered and not adjusted for multiple comparisons.

What to watch next

  • Preregistered replication using real booking choices and specific clinical AI tasks rather than a generic disclosure.
  • Tests of bare versus detailed disclosure, including clinician responsibility, patient choice and locally verified performance.
  • Cross-country and multilingual studies in actual health systems with observed opt-outs, complaints and second-opinion requests.
  • Whether effects persist after patients know the clinician, the care context and the AI system's limits.

Evidence trail

Sources used for this report

Links checked 1 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Society & Media

Can image-aware AI catch fake news before it spreads?

A new peer-reviewed model improved two fixed benchmark tests by combining article text, images and an AI-generated image description. It did not verify claims or face a live, changing news stream.

7 min · 2 sources

Society & Media

Can an inaccurate AI summary change what you remember seeing?

In a US online experiment with 328 analysed participants, accurate recall of a traffic sign fell from 83.6% after a consistent summary to 44.8% after a misleading one. The study isolates a classic misinformation effect; it does not measure real police reports or prove that ordinary model errors cause the same size of harm.

7 min · 2 sources

Society & Media

Did AI-written petitions persuade more people to act?

A peer-reviewed natural experiment covering 1.5 million Change.org petitions found that access to an embedded AI writer changed language and increased similarity, but did not improve the engagement outcomes the researchers measured.

5 min · 2 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.