ANALYSISAcross 24 official German licensing exams, the best model answered 99.31% of first-exam items correctly. On items containing an image, the error rate rose several-fold, against 1.24x for students.
4 min read
ANALYSISA system called Quicker writes clinical guideline recommendations through a GRADE workflow. A Matters Arising and its reply, published the same day, map what still has to be built around it.
3 min read
ANALYSISA blinded comparison of 180 responses from a university telehealth centre in Minas Gerais found AI matched human specialists on medical adequacy and risk, and beat them on comprehensibility.
3 min read
ANALYSISTen studies, wide confidence intervals, and factual error rates of 26 to 36 percent in AI-drafted documentation. The review's own conclusion is that the evidence is preliminary and highly uncertain.
4 min read
ANALYSISPhysicians editing AI drafts wrote better notes than they did unaided, and worse notes than the AI produced alone. Both directions of that result matter.
5 min read
ANALYSISTwenty experts scored three language models on ten common questions. One model cleared the content validity bar on both topics; another scored zero on jet lag.
4 min read
ANALYSISAn EHR-embedded model wrote hospital summaries residents edited less — but with more confabulations. A separate pipeline audited 21,041 trial reports against CONSORT at 91.7% agreement with experts.
5 min read
ANALYSISGoogle's PH-LLM beat sampled human experts on multiple-choice tests but only matched them on real cases. A new reporting checklist published the same month explains why such claims are hard to compare.
4 min read
ANALYSISSingapore General Hospital randomised residents to work with and without an LLM assistant. Documentation time fell by 1.82 minutes, which was not significant. The economic model used the point estimates anyway.
5 min read