NICE reviewed 30 digital mental health tools. The same evidence gaps kept recurring
A cross-sectional analysis of NICE evaluations found 78 supporting studies behind 30 technologies — and consistent holes in comparators, cost of delivery and adverse-event reporting.
Digital mental health technologies have moved from the margins of NHS mental health services toward the centre of them, and England's National Institute for Health and Care Excellence has been the gate they pass through. A study published in JMIR Mental Health asked a narrow, useful question about that process: what does the evidence actually look like in the documents NICE committees were reading [s1]?
The answer is not that the evidence is bad. It is that it is uneven in a patterned way, and the pattern repeats across products.
The method
The authors identified every NICE evaluation relating to digital mental health technologies by reviewing published evaluations on the NICE website [s1]. From each, they pulled out the individual technologies assessed and worked through the committee documentation to find the studies cited in support of each one [s1].
That produced nine NICE evaluations, covering anxiety, depression, psychosis, insomnia, attention deficit hyperactivity disorder and tic disorders [s1]. Those evaluations took in 30 separate digital technologies, and referenced 78 supporting studies between them [s1].
The team then extracted a series of items relating to study quality, and summarised the evidence at two levels: the individual study, and the whole package of studies backing a given technology [s1]. The second level matters more than it sounds. A regulator does not judge a product on its best trial; it judges the assembled case.
What kept being missing
Four gaps recurred across the evidence packages [s1].
The first was effectiveness against relevant comparators. Showing that a digital therapy beats a waiting list is a much weaker claim than showing it beats the care a patient would otherwise receive, and the distinction determines whether adopting the technology improves anything.
The second was outcomes. The authors flagged the use of appropriate outcomes, including health-related quality of life, as a common shortfall [s1]. Health technology assessment in England runs on quality-adjusted life years, which require quality-of-life measurement. A trial that reports only a symptom scale leaves the committee unable to do the arithmetic its own framework demands.
The third was cost — specifically the cost of delivery and the impact on resource use [s1]. Digital tools are often pitched as cheap, but the relevant figure is not the licence fee. It is what it costs to run the thing inside a service: the clinician time for triage and review, the technical support, the referral pathway. That is what determines whether a technology releases capacity or quietly consumes it.
The fourth was adverse events [s1]. Reporting of harms was identified as a gap in the evidence base. Psychological interventions can produce deterioration in a minority of people, and a digital product delivered at scale distributes any such effect widely. If trials do not report harms, the assessment cannot weigh them.
What the authors did not claim
The paper's conclusion is more balanced than the gap list suggests. Some digital mental health technologies have been supported by high-quality studies, the authors note, and evidence for a given technology is likely to be developed across a series of studies rather than resting on a single trial [s1]. The argument is that the recurring gaps need addressing to make a stronger case for adoption, and that developers should plan for them earlier in the product lifecycle [s1].
This is also a study of documents, not of patients. It describes what evidence was submitted and cited; it does not measure whether the technologies work, whether NICE's decisions were right, or whether patients using them fared well or badly. A cross-sectional analysis of nine evaluations cannot answer any of those questions, and does not try to.
The sample is small by design. Nine evaluations and 30 technologies is the whole population of NICE digital mental health assessments, not a subset — but it is still nine committees' worth of decisions, and the patterns identified could shift as the pipeline grows.
Why it matters beyond England
NICE is watched well beyond the UK. Its methods are studied by health technology assessment bodies across Europe, and the EU's joint clinical assessment machinery under the health technology assessment regulation has been building its own approach to exactly this class of product. A finding that comparator choice, quality-of-life measurement, delivery cost and harms reporting are the four places digital mental health evidence tends to thin out is portable, because those four requirements are common to almost every European assessment framework.
For a reader trying to interpret claims about a mental health app, the practical translation is a set of questions rather than a verdict: what was it compared against, was quality of life measured, what does it cost to run rather than to buy, and were harms looked for at all. The study suggests that for a meaningful share of products reaching assessment, at least one of those answers has been missing [s1].
Sources
- [s1] Strength of Evidence to Support Decision-Making on the Use of Digital Mental Health Technologies in NICE Evaluations: Cross-Sectional Analysis of Studies — JMIR Mental Health, published online 7 April 2026. https://doi.org/10.2196/85635
Sources
More on
Sleep got worse during Ramadan and mood mostly didn't — in a cohort of 30 Saudi women
Stress peaked during the fasting month and fell afterwards, but depression and anxiety stayed flat. The consistent finding across every phase was that sleep quality tracked mental health.
Only 5% of Lebanese children who screened positive had ever received care
The barriers parents named most often were cost and the absence of any nearby service — not stigma, which the field has spent years treating as the main obstacle.
A 2967-adolescent cohort looked for bushfire harm at 24 months and did not find it
Australian teenagers harmed by the Black Summer fires showed no elevated depression, anxiety, distress, insomnia or suicidality two years later. The result runs against a systematic review published months earlier.
A ten-society guideline on tapering benzodiazepines, and data on who is starting them
The guidance is built on one negative instruction: do not stop abruptly in patients likely to be physically dependent. Hong Kong records show the sharpest prescribing rise in 18-to-25-year-olds.