Clinical AI Evidence 2026

Five peer-reviewed findings on ambient AI scribes, clinician burnout and diagnostic reasoning — with the exact figures, conditions and limits of each study.

Respocare Connect AI Team — Respocare (PTY) Ltd · Reg. 2018/411829/07 · Practice No. 9990900010775614. Licensed healthcare practice operating since 2018.

· 10 min read

The Clinical AI Evidence: Five Findings, and What Each One Actually Says

A reference review of the peer-reviewed evidence on ambient scribes, documentation burden and diagnostic reasoning — including the conditions each result depends on.

Last reviewed: 9 September 2026

Artificial intelligence attracts extraordinary claims. This page collects five of the findings most often quoted in clinical AI conversations, states each one with its exact figures and study conditions, and says plainly what it does and does not establish.

None of these studies proves that any given AI system is clinically safe, or that a result achieved in one setting transfers to another. What they show, taken together, is that the question has changed. The field is no longer asking whether generative AI has a meaningful use case in healthcare. It is asking how to integrate it responsibly — which is a harder question, and a more interesting one.

1. Burnout fell from 51.9% to 38.8% after thirty days of ambient scribe use

A quality improvement study of 263 physicians and advance practice practitioners across six health care systems found that after 30 days with an ambient AI scribe, burnout among those working in ambulatory clinics decreased significantly from 51.9% to 38.8%, alongside significant improvements in cognitive task load, time spent documenting after hours, focused attention on patients, and urgent access to care. The 51.9% to 38.8% shift was measured among the 186 participants included in the burnout models — a difference of 13.1 percentage points, corresponding to an adjusted odds ratio of 0.26.

The conditions that matter. The study ran between February and October 2024 across six US health systems deploying a single commercial ambient scribe product. There was no control arm, and the endpoint was self-reported. It is a quality improvement study, not a randomised trial.

What it establishes. Not that an AI scribe cures burnout. Something more interesting: changing how clinical documentation happens may change the experience of practising medicine. That is a larger claim than faster typing, and it is the reason documentation became the entry point for clinical AI rather than diagnosis.

Olson KD, Meeker D, Troup M, et al. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout. JAMA Network Open. 2025;8(10):e2534976.

2. The documentation burden the first finding is measured against

The burnout result only means something set against the baseline. Physicians have been reported to spend nearly half the office day on the electronic health record and desk work, and roughly a quarter of it face-to-face with patients — with a further one to two hours of EHR work after hours.

Consider the human economics of that. We put people through years of training, hand them some of the most consequential decisions in society, and then allow a large share of their working day to be consumed by moving information from one place to another.

This is why the opportunity should not be measured only in minutes saved per note. The better question is what medicine does with attention returned to it — and that is not a question a time-and-motion study can answer.

Sinsky C, Colligan L, Li L, et al. Allocation of Physician Time in Ambulatory Practice: A Time and Motion Study in 4 Specialties. Annals of Internal Medicine. 2016;165(11):753-760. doi:10.7326/M16-0961

3. Ambient scribes are now the most widely implemented generative-AI solution in healthcare

A January 2026 JAMA Network Open commentary on return on investment opens by stating that ambient AI scribes have rapidly become the most widely implemented generative AI solution in health care, and notes that they remain costly, with most health systems paying subscription fees of $200 to $600 per clinician per month. The commentary accompanies one of the first quantitative evaluations of the ROI question, examining physician financial productivity across more than 1.2 million ambulatory encounters at a single academic site.

What it establishes. That adoption has outrun evaluation, in an industry where implementation traditionally takes years. The scribe solved something structural about how new technology enters medicine: it does not begin by replacing clinical judgement, it begins by removing the work surrounding it.

The same commentary makes a second observation worth sitting with — that institutions investing in the infrastructure, governance and training for responsible scribe integration are simultaneously preparing for what comes next: clinical summarisation, automated coding and real-time decision support. The scribe is the doorway. It is not the room.

Shah SJ, Garcia P. Ambient AI Scribes — What Is the Return on Investment? JAMA Network Open. 2026;9(1):e2553238.

4. A model scored 92%, and giving physicians access to it changed almost nothing

This is the finding that should govern how anyone builds clinical AI.

In a randomised clinical trial, fifty physicians — 26 attendings and 24 residents — worked through clinical vignettes either with GPT-4 alongside conventional resources, or with conventional resources alone. The median diagnostic reasoning score per case was 76% for the group with the model and 74% for the group without, an adjusted difference of 2 percentage points that was not statistically significant. Across three runs of the model working alone, the median score per case was 92%.

What it establishes. Not that doctors are unnecessary. Close to the opposite. The authors' own conclusion is that access alone to language models will not improve overall physician diagnostic reasoning in practice, and that the gap indicates a need for technology and workforce development to realise the potential of physician–AI collaboration.

Access to intelligence is not the same as integrating intelligence. A model can become dramatically more capable and still fail to improve the workflow if the interaction between clinician and machine has not been designed. Interface, context, training, workflow and governance decide the outcome — not raw capability.

The future cannot simply be a chatbot for every doctor.

Worth noting. A later randomised trial of a purpose-built collaborative system, using identical vignettes and scoring, did significantly improve clinician performance — by 9.9% where the AI gave a first opinion and 6.8% where it gave a second — while its AI-alone benchmark sat at a median of 89.5%, close to the earlier 92%. The model did not get better. The workflow around it did. That is the whole argument.

Goh E, Gallo R, Hom J, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open. 2024;7(10):e2440969.

Everett SS, Bunning BJ, Jain P, et al. From tool to teammate in a randomized controlled trial of clinician-AI collaborative workflows for diagnosis. npj Digital Medicine. 2026. doi:10.1038/s41746-026-02545-1

5. Ambient AI increased eye contact between physician and patient

A prospective within-clinician time-motion study at a large academic medical centre in Singapore, running from December 2024 to May 2025, found ambient scribe use associated with a 15.0% reduction in documentation time per consultation, a 10.6% increase in the proportion of consultation time involving eye contact — from a mean of 69.6% to 77.1% — and no significant change in consultation duration or total cycle time per patient. Of 39 patients surveyed, 27 agreed their physician focused on them more during the consultation, and none expressed discomfort with the technology.

The conditions that matter. Nine clinicians, 169 consultations, five trained observers, an in-house multilingual tool. This is a small, single-site study of experienced users. It should be read as a signal, not a benchmark.

What it establishes. That the measurable effect of documentation technology can appear somewhere other than the clock. Perhaps the best measure of clinical technology is not how much attention the machine receives, but how much attention it gives back to the people in the room.

Impact of an Ambient AI Scribe Among Clinicians and Patients: Real-World Prospective Observational Time-Motion Study. JMIR Medical Informatics. 2026;14:e85580. doi:10.2196/85580

What this evidence is not

Three of these five findings concern ambient scribes as a product category. They are not Respocare Connect AI results, and nothing on this page is a claim about our platform's performance.

Our production Medical AI Scribe today is a post-encounter dictation model. Ambient capture is in development and we will say so here when it is live. Respocare Connect AI is a clinical assistant and documentation tool. It is not a medical device, it is not diagnostic, and every output requires independent clinician review before it enters a patient record.

We would rather under-describe our own system than borrow somebody else's evidence for it.

Where the evidence points next

The scribe is the first chapter rather than the destination. It gave healthcare an understandable way to begin: listen, structure the encounter, draft something useful, return it to the clinician for approval.

Once intelligence is safely present in the workflow, the harder questions arrive.

Can the system retrieve the relevant history? Reason across encounters rather than within one? Detect a trend that sits across several documents? Separate what is documented from what is inferred? Hold a safety-critical fact intact as a record grows and fragments? Perform controlled actions under governance? And can it recognise the moment to do none of those things and hand control back?

At that point the conversation stops being about scribes and becomes a conversation about agency — where the fourth finding on this page stops being an interesting result and becomes a design constraint. A more capable model does not produce a better clinical outcome on its own. The architecture around it decides that.

Respocare Connect AI is built to that principle. Our agentic clinical assistant, North Star, is designed to retrieve, reason, act, refuse and escalate within a governed loop, with deterministic checks around the model and a draft — never a decision — as the output.

The AI drafts, the record verifies, the clinician decides.

Frequently asked questions

Does AI reduce physician burnout?

One quality improvement study of 263 clinicians across six health systems found burnout fell from 51.9% to 38.8% after 30 days of ambient AI scribe use. The study had no control arm and used self-reported measures, so it shows an association rather than proof of cause.

Do AI scribes make consultations longer or shorter?

A 2026 time-motion study found documentation time per consultation fell by 15.0% while consultation duration itself did not change significantly. The study observed nine clinicians across 169 consultations at a single site.

Are AI models better at diagnosis than doctors?

In one randomised trial, a model working alone scored a median of 92% on diagnostic reasoning cases, while physicians given access to that same model scored 76% against 74% for physicians without it. The gap between the model's standalone score and the physician-plus-model score is the central finding, and it points to workflow design rather than model capability.

Is Respocare Connect AI clinically validated?

No. Respocare Connect AI is built to a proof-of-promise engineering standard, and independent clinical validation is ongoing rather than complete. We do not describe the platform as clinically validated and will not until that work is finished and reviewed.

Is Respocare Connect AI a medical device?

No. It is a clinical assistant and documentation tool. It is not diagnostic, it is not a substitute for clinical judgement, and all output requires independent clinician review.

Built by Respocare (PTY) Ltd — a licensed South African healthcare practice operating since 2018.