An Israeli AI cut antibiotic mismatch by 30%. A third of doctors ignored it anyway
Maccabi Healthcare Services studied 626 of its own physicians to work out who follows algorithmic prescribing advice — and found the pattern was about practice structure, not just attitude.
The Israeli health fund Maccabi Healthcare Services introduced a machine-learning decision support system for urinary tract infection prescribing in 2021 [s1]. The tool, called UTI Smart-Set, reduced antibiotic mismatch — cases where the pathogen recovered on culture turned out to be resistant to the empirically prescribed antibiotic — by roughly 30% [s1].
Around 33% of physicians did not follow its recommendations [s1]. A retrospective cohort study published on 11 June set out to characterise which ones.
The design
Researchers used Maccabi's own data covering UTI encounters between September 2023 and March 2024, analysing 626 physicians across 15,033 encounters [s1]. They examined correlations between physician characteristics and whether the physician implemented the system's recommendations, adjusting for patient- and encounter-level variables [s1].
Studying adoption inside a single integrated health fund has a specific advantage: every physician in the sample was working with the same tool, the same electronic record, the same formulary and the same patient population structure. Variation in uptake cannot be attributed to differences in the technology itself.
What predicted following the recommendation
Three characteristics were associated with higher odds of implementing the system's advice: seeing a younger patient population (odds ratio 0.952 per year of mean patient age, 95% CI 0.922–0.983), diagnosing more UTIs (OR 1.021 per case, 95% CI 1.007–1.035), and working within a group practice rather than solo (OR 1.542, 95% CI 1.02–2.333) [s1].
The group-practice finding is the most actionable of the three. Adoption of a clinical tool is often framed as an individual disposition — some doctors trust algorithms, some do not. An association with practice structure suggests something more amenable to intervention: shared norms, peer discussion of cases, or simply the local diffusion of a habit among colleagues who see each other.
The volume finding runs in a direction that may be counterintuitive. Physicians who diagnosed more UTIs were more likely to follow the recommendation, not less — the pattern that "I've seen a thousand of these, I don't need the machine" would predict the opposite. One reading is that repeated exposure to the tool in the condition it was built for produces familiarity and calibration.
What predicted ignoring it
Three characteristics were associated with lower odds of implementation: older physicians (OR 1.034 per year, 95% CI 1.012–1.056), practising in the Arabic sector (OR 3.474, 95% CI 1.709–7.062), and carrying a higher patient volume overall (OR 1.027 per 100 patients, 95% CI 1.003–1.052) [s1].
The Arabic-sector finding has the largest effect size in the study and the widest confidence interval, which reflects a smaller subgroup. The paper reports it as an association and does not establish what drives it. Language of the interface, the specific resistance patterns in the populations served, practice conditions, clinic infrastructure, and consultation time pressure are all candidates, and the study design cannot distinguish between them. Reading an association like this as being about the physicians rather than about the circumstances they work in would go well beyond what the data supports.
The overall patient-volume finding pulls in the opposite direction from the UTI-specific volume finding: more UTIs meant more adoption, more patients overall meant less. That combination is consistent with time pressure being the mechanism — a busy general list leaves less room to engage with a prompt — but the study does not test it.
Why this matters beyond one condition
Antibiotic mismatch in UTI is a good test case for clinical AI because the ground truth arrives on its own. A urine culture returns days later and settles definitively whether the empiric choice would have worked. Most clinical decision support operates without that kind of clean feedback, which is why demonstrated benefit is rarer than deployment.
The 30% mismatch reduction is stated in this paper as background — the previously established effect of the system, not a finding of this study [s1]. What this study establishes is the implementation ceiling: a tool with a demonstrated benefit reached about two-thirds of prescribing decisions in the period studied.
The limits
This is retrospective and observational, covering seven months at one health fund. It identifies correlates of adherence, not causes, and it cannot show whether changing any of the identified characteristics would change behaviour.
It also does not report outcomes by adherence — whether patients of physicians who followed the recommendations did better than patients of those who did not. The 30% mismatch reduction applies to the system, not to this comparison.
And "did not follow the recommendation" is not the same as "was wrong". A clinician who overrides a decision support prompt may be responding to information the model does not have — allergy history, prior culture results, pregnancy, a detail from the conversation. The study measures implementation, and the authors frame the goal as improving integration [s1] rather than as eliminating override.
What to watch next
Whether Maccabi acts on the practice-structure finding. If group practice really does drive adoption through shared norms, then interventions aimed at solo and high-volume practices — peer feedback, audit-and-feedback loops, workflow redesign — would be testable, and the health fund has the data infrastructure to test them.
Sources
- Physician adoption patterns of AI-driven clinical decision support systems in urinary tract infection management — Scientific Reports, 11 June 2026 (primary)
Sources
- Physician adoption patterns of AI-driven clinical decision support systems in urinary tract infection management — Scientific Reports , June 11, 2026
More on
An Israeli HMO biobank sequenced 1,038 patients to hunt for deafness genes
Linking exome data to electronic medical records solved 15% of unexplained hearing-loss cases and flagged new candidate genes — while showing what the records could not supply.
The ACP has issued ethical guideposts for AI at the bedside. There are three.
Relationality, self-governance, competence. The position paper's starting premise is that consensus on privacy, disclosure and fairness has not been reached, and clinicians need guidance anyway.
Sepsis AI reaches an AUROC of 0.88 and a positive predictive value of 34.2%
A network meta-analysis of 53 studies and more than 7 million admissions finds machine learning out-discriminates traditional sepsis scores — and would raise roughly two false alarms for every real one.
UAE hospitals are using AI to flag sepsis six hours before doctors would catch it
The tool reads vital signs, lab results, and history the way clinicians already do — just faster and continuously. Its accuracy tops out around 90%, which means it's also still wrong a meaningful share of the time.