ANALYSISA Danish stepped-wedge trial gave endoscopists automated feedback on their technique after every procedure. Adenoma detection rose from 43.4% to 48.6% — a different tool from real-time polyp AI.
3 min read
ANALYSISTrained on 1.7 million ECGs paired with clinicians' report text, ECG-CLIP reached the same accuracy as the best comparator with about 90% less training data — a bid at the field's labelling bottleneck.
3 min read
In 6,772 users of a US health system's AI triage tool, people engaged about twice as often when the AI matched what they already planned — raising a hard question about what the tools are steering.
3 min read
A review of 41 studies found that the great majority of AI medication-adherence prediction models carried high risk of bias, and that fancier algorithms did not reliably predict better.
3 min read
ANALYSISAcross 24 official German licensing exams, the best model answered 99.31% of first-exam items correctly. On items containing an image, the error rate rose several-fold, against 1.24x for students.
4 min read
ANALYSISA researcher who expected the evidence base to be thin says even she was surprised by how thin. Most cleared devices never appear in a registered clinical trial at all.
3 min read
A new discussion paper proposes a two-axis risk framework and a physician-training analogy for evaluating generative AI devices. It is not a rule, and the agency is asking what one should look like.
4 min read
ANALYSISA Nature Reviews Drug Discovery audit says the problem isn't the models. They were built to be validated rather than used, and benchmarked against the wrong thing.
4 min read
ANALYSIS42 orthopedic physicians diagnosed 40 rare diseases twice, once alone and once after seeing AI suggestions. Accuracy jumped 20 to 26 points, but the same-day design leaves memory unaccounted for.
3 min read
ANALYSISA Lancet Digital Health scoping review found the field's fairness metrics fragmented and rarely clinically validated. A second review found that most studies don't measure fairness at all.
4 min read
A review of Paige Prostate Detect and Ibex Prostate Detect finds the tools mainly help less-specialized pathologists, and warns performance shifts when a tool meets a new population.
3 min read
EXPLAINERTwo 2026 reviews make the case that human-derived models can replace animal testing. A third paper argues the field is checking the biology and the algorithm separately, and calling that assurance.
4 min read
A paper in the journal Resuscitation lays out how a large language model could help dispatchers recognize cardiac arrest and coach CPR in real time — while acknowledging the concept still needs to prove it saves lives.
3 min read
ANALYSISA small proof-of-concept study pitted several AI systems against gastroenterologists and emergency physicians on cholangitis exam questions. The gap between the best and worst AI performers was enormous.
2 min read
The agency granted its first authorization for a software-aided adjunctive diagnostic device in wound assessment, a category built for tools that analyze a wound optically.
2 min read
ANALYSISA study running 5.3 million evaluations through nine large language models found eligibility judgments were largely stable across identity labels — except when a patient vignette mentioned homelessness.
3 min read
ANALYSISDxDirector-7B beat human physicians on a benchmark of complex diagnostic cases while requesting far fewer tests. Its own authors say it isn't ready for high-risk or emergency cases.
3 min read
ANALYSISA secondary analysis of the VITAL-AF trial reports a screening benefit only in the top risk decile — with a confidence interval whose lower bound sits at 0.01.
4 min read
GE's Critical Care Suite gains an algorithm that flags misplaced enteric tubes on a chest X-ray — a complication that reviews of the practice have linked to respiratory harm and, in some cases, death.
2 min read
The clearance is the third in a growing family of machine-learning notification algorithms built to spot serious cardiac conditions from a routine 12-lead ECG, a test most patients already get.
2 min read
ANALYSISA new audit of China's regulatory record finds a market concentrated in radiology, dominated by deep learning, and clustered in four cities. It also finds the approval curve flattening.
3 min read
ANALYSISA 1,298-person randomized study found people using chatbots to work through medical scenarios did no better than people without them — even though the same chatbots, tested alone, got the right answer most of the time.
3 min read
ANALYSISResearchers ran 960 responses through ChatGPT Health using clinician-written vignettes. Failures clustered at both extremes, and crisis safeguards activated unpredictably.
4 min read
ANALYSISA February theme issue documents faster notes, happier clinicians and enterprise rollouts to thousands. It also contains an editorial asking the question none of the studies answer.
5 min read
ANALYSISTen studies, wide confidence intervals, and factual error rates of 26 to 36 percent in AI-drafted documentation. The review's own conclusion is that the evidence is preliminary and highly uncertain.
4 min read
WHAT THE STUDY ACTUALLY SAYSAlphaGenome takes a million base pairs of DNA and predicts what the sequence does. In Nature on January 28, it matched or beat the best existing models in 25 of 26 variant-effect evaluations.
4 min read
ANALYSISClinicians rated empathy at 4.6 out of 5 and quality of information at 2.7. Every chatbot produced at least one piece of guidance judged inappropriate, overstated or inaccurate.
4 min read
ANALYSISThe EAGLE trial ran colonoscopy AI off-site over a network and aimed it at the lesions that matter. It reports a threefold gain in serrated lesions — in a field whose main US guideline recommends nothing.
5 min read
ANALYSISOne agency asked how to measure AI devices after they are deployed. Another asked how to speed adoption up. The offices that would answer the first question are losing staff.
4 min read
ANALYSISThe headline finding is noninferiority. The more useful finding is what the shared denominator was — and how wide a gap the trial was designed to tolerate.
3 min read
ANALYSISEngland ran the closest thing to a national test of clinical AI anyone has published. The software sites gained more, but the comparison sites were improving on their own, and that gap is the whole result.
4 min read
ANALYSISThe WISeR model starts on 1 January in six states. The contractors running it are paid out of the savings their determinations produce — the design feature doctors keep pointing at.
4 min read
Researchers showed open-source protein design software could produce variants of proteins of concern that biosecurity screening tools missed, then wrote and deployed patches before publishing.
3 min read
ANALYSISA census of 691 AI-enabled devices cleared through 2023 found demographic data missing almost everywhere, three summaries reporting patient outcomes, and 113 recalls driven mostly by software.
4 min read
ANALYSISEyeFM was tested as an assistant to 16 ophthalmologists screening 668 high-risk patients in China. Patients in the AI arm also followed referral advice more often.
4 min read
ANALYSISGoogle's PH-LLM beat sampled human experts on multiple-choice tests but only matched them on real cases. A new reporting checklist published the same month explains why such claims are hard to compare.
4 min read
ANALYSISEchoNext scored 77.3% accuracy on a 150-ECG set where 13 cardiologists averaged 64.0%. Its authors released the model weights and a 100,000-ECG labelled dataset alongside the paper.
6 min read
ANALYSISA draft-reporting model cut documentation time 15.5% across 23,960 radiographs without changing report quality. A separate reader study found AI assistance raised prostate MRI accuracy by 3.3 percentage points.
4 min read
ANALYSISRentosertib was safe over 12 weeks in 71 patients with idiopathic pulmonary fibrosis. A lung-function signal appeared at the highest dose, in a trial not designed to prove efficacy.
4 min read