WHAT THE STUDY ACTUALLY SAYS

The NICE fever traffic light: what it is, and how well it works

The system sorts feverish under-5s into red, amber and green. Validation studies find it misses serious illness and over-refers the well — a prompt, not a reliable test.

The NICE traffic-light system is a chart that sorts feverish children under five into red (high-risk), amber (intermediate) and green (low-risk) groups by their colour, alertness, breathing, hydration and — in young infants — temperature, to flag those who might have a serious illness [s1]. It is widely used, but when researchers have checked it against what actually happened to children, it has performed poorly, missing many who turned out to be seriously ill while flagging a large majority who were not — one UK validation concluded it is "not suitable ... as a clinical tool in general practice" [s2].

What the tool actually says

NICE guideline NG143, first published in 2019 and last updated in November 2021, groups warning features into three colours [s1]. Red features include pale, mottled, ashen or blue skin; no response to social cues; appearing ill to a healthcare professional; not waking, or not staying awake when roused; a weak, high-pitched or continuous cry; grunting; a respiratory rate above 60 breaths a minute; and reduced skin turgor [s1]. Amber features are milder versions — pallor reported by a parent, poor feeding in infants, reduced urine output, dry mucous membranes, or a capillary refill time of three seconds or longer [s1]. A child with none of the red or amber signs, who is a normal colour, responds normally, and is well hydrated, is classed green [s1].

Temperature carries specific weight only at the extremes of infancy: NICE places any child younger than three months with a temperature of 38°C or higher in the high-risk group, and any child aged three to six months with a temperature of 39°C or higher in at least the intermediate group [s1]. Beyond those ages, the guideline is explicit that the height of the fever alone should not be used to identify serious illness [s1].

How well it performs

The tool's face validity is not the same as accuracy, and two validation studies have tested it directly. In a retrospective cohort of 6,703 acutely unwell children under five in English and Welsh general practice, linked to hospital records, the system classed 2,116 (31.6%) as red, 4,204 (62.7%) as amber and just 383 (5.7%) as green [s2]. Serious illness was rare — 139 children (2.1%) were admitted within seven days, of whom 17 (12.2%, or 0.3% of the whole cohort) had a serious illness [s2]. Against that outcome, the red category had a sensitivity of only 58.8% (95% confidence interval 32.9% to 81.6%) and a specificity of 68.5% (67.4% to 69.6%) — meaning it missed roughly four in ten seriously ill children while still labelling a third of the well ones red [s2].

Widening the net to red-or-amber caught everyone who was seriously ill (sensitivity 100%, 95% confidence interval 80.5% to 100%) but at the cost of near-useless specificity of 5.7% (5.2% to 6.3%) — almost every child, sick or not, screens positive [s2]. The authors' conclusion was blunt: the system did not accurately separate children with serious illness from those who could be managed at home, and is not suitable as a general-practice tool [s2].

An earlier study of the individual red signs found the same weakness from the other direction. Validating the 16 most severe NICE features across seven primary-care and emergency settings in 6,260 children, it found that only four features — not waking or staying awake, reduced skin turgor, a non-blanching rash, and focal neurological signs — meaningfully raised the probability of serious infection in more than one dataset, and even the presence of three or more red features fell short of a strong rule-in signal [s3]. The alarming signs, in other words, are better at raising suspicion than at confirming it.

Why this matters

None of this means fever in children should be ignored; it means a widely deployed checklist is a rough prompt rather than a precise test, and its numbers should be read with that limit in mind [s2][s3]. A tool that flags most children amber will, by design, generate more referrals and investigations than serious illness alone would warrant — a known cost of ruling out rare, dangerous infections early. The research point that recurs is that alarm-symptom checklists deserve validation against real outcomes before they are built into routine practice, not after [s3].

How to read this

The traffic-light system is best understood as a structured way to notice red flags, whose real-world accuracy is modest, rather than as a rule that reliably tells serious from self-limiting illness [s2][s3]. It sits alongside two related questions parents ask about fever, covered separately: whether and how to bring a temperature down in treating fever, and the specific worry of a convulsion in febrile seizures.

This article is informational and not medical advice; a child who is unwell or causing concern should be assessed by a clinician without delay.

Sources

  1. Fever in under 5s: assessment and initial management (NG143) — National Institute for Health and Care Excellence (NICE) , November 26, 2021
  2. Accuracy of the NICE traffic light system in children presenting to general practice: a retrospective cohort study — British Journal of General Practice , March 24, 2022
  3. The Predictive Value of the NICE 'Red Traffic Lights' in Acutely Ill Children — PLoS ONE , March 14, 2014

More on

Related coverage