Graders spotted the chatbot in Brazil's telehealth test, and rated it clearer anyway
A blinded comparison of 180 responses from a university telehealth centre in Minas Gerais found AI matched human specialists on medical adequacy and risk, and beat them on comprehensibility.
Brazil's Telehealth Brazil Program routes clinical questions from primary care workers to specialists at university telehealth centres, which answer them in writing. A study published in Telemedicine and e-Health on 25 July took 180 of those exchanges from one such centre and asked whether a large language model could have written the answers instead [s1].
The design
The comparison drew on records from the Telehealth Center of the UFMG Faculty of Medicine (NUTEL FM-UFMG), a member of the national program, covering teleconsultations submitted between January 2020 and May 2024 across three specialties: cardiology, endocrinology, and obstetrics and gynaecology [s1].
For each question, the study set a real human specialist response against an AI-generated one, then had them assessed blind against four quality criteria — medical adequacy, conciseness, coherence and comprehensibility — plus risk potential, accuracy of authorship identification, and whether the inquiry was resolved [s1]. Analysis used the Shapiro-Wilk test, Kruskal-Wallis test, Nemenyi multiple comparison test, chi-square test and Fisher's exact test [s1].
What the graders found
Across all three specialties, one difference reached statistical significance: comprehensibility, where AI mean scores exceeded human scores [s1]. On the remaining quality criteria, as well as on risk potential and inquiry resolution, no significant differences emerged — though the study notes AI scored higher than humans on those measures too, without the gap reaching significance [s1].
Within individual specialties, significant differences favouring AI appeared in endocrinology across all criteria except conciseness, and in cardiology for conciseness [s1].
The finding that complicates the rest is authorship identification. Across all specialties, and individually within endocrinology and obstetrics/gynaecology, graders' accuracy at telling human from AI responses was statistically significant [s1] — meaning the blinding did not hold. Assessors could distinguish the two better than chance.
That matters for interpreting everything else. When evaluators can identify which response came from a machine, their quality ratings are no longer independent of that knowledge, and the direction of any resulting bias is not something this design can measure. The study reports the identification result as one of its outcomes rather than treating it as a threat to the others, but a reader should hold the quality findings more loosely because of it.
What "no significant difference" is and is not
The core result here is a set of null findings on medical adequacy, coherence, risk potential and resolution. A null finding is not equivalence. With 180 responses split across three specialties, the study is not powered to rule out modest differences in clinical quality, and it does not report an equivalence margin. "We could not detect a difference" and "there is no difference" are different statements, and only the first is supported.
The one positive finding — better comprehensibility — is also the least clinically loaded. Readability is a genuine property of a good specialist answer, particularly one that a nurse or family physician in a remote municipality has to act on. It is not the same as being right.
Why it is being asked in Brazil specifically
The study frames telehealth as a strategic component of primary health care that has advanced in Brazil through the National Telehealth Program, and positions AI as a way to enhance its benefits [s1]. The operational appeal is capacity: teleconsultation services depend on specialist time, and specialist time is the scarce input in a system serving a population distributed across a very large territory. A tool that drafts responses for specialist review changes the economics of that service in a way that adding specialists does not.
The study's own conclusion is appropriately narrow — that despite existing limitations, AI demonstrates substantial potential as a support tool for teleconsultation services [s1]. Support tool, not replacement, is the operative framing, and it is the framing the evidence here actually supports.
What the study does not cover
Several things. It measures response quality as judged by evaluators, not outcomes in the patients whose care prompted the questions. It covers three specialties at one centre. The source questions span January 2020 to May 2024 [s1], a period during which both clinical practice and the models themselves changed considerably, and the available reporting does not describe how question difficulty was distributed across that window. And it does not report what happens to quality when a specialist edits an AI draft rather than choosing between the two — which is the workflow any real deployment would use.
What to watch
Whether Brazilian telehealth centres move from evaluation studies to supervised deployment, and whether anyone measures the thing this study could not: whether questions answered with AI assistance lead to different clinical decisions in the primary care units that asked them.
Sources
- Performance of Large Language Model-Based Chatbots in Primary Health Care Teleconsultations: A Comparison Between Human and Artificial Intelligence-Generated Responses — Telemedicine and e-Health, 25 July 2026
Sources
- Performance of Large Language Model-Based Chatbots in Primary Health Care Teleconsultations: A Comparison Between Human and Artificial Intelligence-Generated Responses — Telemedicine and e-Health , July 25, 2026
Audio-only telehealth visits rate worse, and Black patients get more of them
Across 90,670 virtual encounters at one US health system, phone visits scored lower on likelihood to recommend. A separate VA cohort of 1.2 million found telehealth cut same-day mental health access sharply.
Models beat German medical students on text — and fell apart on the picture questions
Across 24 official German licensing exams, the best model answered 99.31% of first-exam items correctly. On items containing an image, the error rate rose several-fold, against 1.24x for students.
An AI drafts guideline recommendations. The argument now is about the guardrails.
A system called Quicker writes clinical guideline recommendations through a GRADE workflow. A Matters Arising and its reply, published the same day, map what still has to be built around it.
An Israeli AI cut antibiotic mismatch by 30%. A third of doctors ignored it anyway
Maccabi Healthcare Services studied 626 of its own physicians to work out who follows algorithmic prescribing advice — and found the pattern was about practice structure, not just attitude.