A British Hospital Missed a Killer for Two Years. The Recommendations Are Mostly Hardware.

I'm traveling internationally for ten days and I turned on one of the British news channels. There's a major story about NHS England. First, the awful reality that a nurse would willfully kill babies in the NICU. Then, the second (and more believable) reality that bad management systems (and poor leadership) made the situation worse.

So what happened?

On July 11th, 2016, the Countess of Chester Hospital put deaths on its neonatal unit onto a “risk register” for the first time. By then, eight babies had died in 2015, and five more had died in the first six months of 2016.

The entry was titled “Apparent Increased Mortality.” It was scored 15 out of a possible 25.

A second entry went on the “Executive Risk Register” the same day:

“Potential damage to Reputation of Neonatal Service and Wider Trust due to Apparent Increased Mortality within the Neonatal Unit.”

That one was scored 20.

A risk score is supposed to be arithmetic: likelihood times consequence. In practice, it is also a prioritization of the queue of possible issues, because the high numbers get the attention and the money. On paper, the Countess of Chester put potential damage to its reputation at 20 and increased mortality at 15.

Ruth Millward, the hospital's Head of Risk and Patient Safety, told the newly-released Thirlwall Inquiry the reputational risk was scored higher because the hospital's ability to reduce it was outside its control.

The response from Lady Justice Kathryn Thirlwall, the Court of Appeal judge who chaired the inquiry, is one plain sentence: “The deaths were not apparent. They were real.”

“The deaths were not apparent. They were real.”

Was the increase in deaths statistically significant? That sounds cold, as I type that. But instead of editing, I'll point out that each death in this case IS significant. But should the number of deaths alone triggered an investigation to look for a “special cause” as it would be called in the Process Behavior Chart methodology.

I had to look it up, but in the NHS, a risk register is a governance document rather than an internal memo. Organizations log what could go wrong, score it on a matrix of likelihood times consequence, name an owner, and review it on a schedule. High scores escalate to the executive register, which the board sees. It's a cousin of FMEA, with one difference that matters here: FMEA also scores detection, asking how likely you are to catch the failure in time. The NHS matrix has no such field.

I don't think anyone sat in a room and decided reputation mattered more than babies. But the scoring is a record of what the organization was actually paying attention to, written down by the people doing the paying attention, and it lines up with almost everything else in the report.

At worst, those numbers are a bad look to the public.

Thirlwall's diagnosis is focused on management and governance. The proposed countermeasures are mostly cameras, locks, protocols, and data returns. That gap is what I want to work through, because almost no organization reading this will ever have a case like Letby, and most of them have a version of the gap.

A note on the convictions, and where just culture stops

The nurse, Lucy Letby was convicted of murdering seven babies and attempting to murder seven others. Those convictions currently stand. An application to the Criminal Cases Review Commission, the independent body that can send a case back to the Court of Appeal, was made in February 2025 and a review is underway.

Thirlwall says at 1.3 of the summary report that her focus was on her terms of reference, the written scope a UK inquiry is given by the government that sets it up, and “not on the guilt of Lucy Letby or on her convictions.” She also refused an application to pause the inquiry pending that review. I'm taking the same position. Nothing below depends on how that appeal turns out, because the management failures the report describes would have been failures whatever the cause of the deaths turned out to be.

I also want to be clear about where systems thinking applies here.

Every version of a just culture framework I know of starts by asking whether the harm was intended. David Marx's model, James Reason's incident decision tree, and the NHS just culture guide built on Reason's work all begin there. NHS England has since retired that guide and replaced it with the “Being fair” tool, which under the Patient Safety Incident Response Framework is used only where there is reason to believe an individual's deliberately malicious, negligent, or incompetent actions contributed to an incident.

If the harm was intended, a just culture analysis does not treat it the same as an honest error or a coaching problem. It calls for a different response, starting with safeguarding and the police. That does not put the system beyond examination. It moves the question from why one person acted to what let the harm continue and why the earlier signals produced nothing.

We shouldn't blame individuals for systemic errors, but we shouldn't use the system as an excuse to avoid accountablity for willfull acts that intended to cause harm.

What is left over, after you set that aside, is a hospital that received a signal and did not act on it for roughly two years. That part is ordinary. That part happens in organizations that have never had anything like this happen.

The Thirlwall Inquiry found the signal was in the data

The report publishes the year-by-year numbers. Deaths on the neonatal unit ran 1, 3, 3, 2, 3 from 2010 through 2014, against annual admissions between 422 and 557. In 2015 there were eight deaths out of 468 admissions.

I put those numbers on a process behavior chart, using only 2010 through 2014 to calculate the limits.

Process behavior chart of Countess of Chester neonatal deaths 2010-2016. Average 2.4, upper limit 5.1, 2015 spikes to 8.

The baseline average (from the data we have) is 2.4 deaths a year. The upper limit is 5.1. That year (2015) with eight deaths is above it.

On an exploratory chart like this one, that's a signal in the technical sense: the previous five years don't explain the cause of the 2015 result. It just tells us, in pretty strong terms, that something changed in the system. It could have been a rash of medication errors. Or, in this case, a killer nurse.

A simple Poisson calculation puts the chance of eight or more deaths at about 1 in 300. Professor Sir David Spiegelhalter, who gave evidence to the inquiry, used a different analysis and assessed it at around 0.008, closer to 1 in 125. Either number was enough, in his words, “to trigger an alert signal, someone should look at this locally.”

In other words, it's VERY unlikely that an unchanged system would have produced eight deaths.

He was also clear about what a chart cannot do. A statistical monitoring system, he told the inquiry, “cannot say why something has happened … a signal can only indicate that someone should look carefully at what is going on.” Professor Marian Knight of MBRRACE-UK put the same point differently: “We have to have both the statistics and the stories.”

I totally agree and I try to teach the same things in my book, Measures of Success.

The small-numbers caution is real and I want to state it clearly. With an average under three, a chart like this is fragile. One year with two extra deaths from unrelated causes moves the point a long way. And a signal on a chart tells you the system changed. It tells you nothing about why.

Three of the eight 2015 deaths were not charged, two from severe congenital abnormalities and one from prematurity with sepsis. If the deaths covered by Letby's murder convictions are removed, the report says, 2015 and 2016 would each show three deaths, consistent with prior years.

Another fine statistical point… as taught by the statistician Dr. Don Wheeler, we can calculate process behavior limits with as few as four data points. But the limits are to be considered provisional — it would be better statistically to have, say, 12 or 15 or monthly data points for those calculations.

A cross-hospital comparison for a single year is more robust than reading a noisy series of six numbers, and it worked. Thirlwall writes at 4.15 that the alert signal “is clearly found in the crude and stabilized and adjusted mortality rates for babies born at the Countess in 2015, as recorded in MBRRACE-UK.”

The problem was time delays. The MBRRACE-UK analysis for 2015 was not reported until June 2017. The 2016 figures came in June 2018. Thirlwall's phrasing:

“The utility of the alert signal provided by the data was lost to time.”

The monitoring system produced a warning that warranted investigation and delivered it after the police had already been called. The trouble was not what it measured. It was how long the result took to reach anyone who could act. The opposite of a real-time metric.

Some of that has been fixed. The MBRRACE cycle is down from 18 months to 14 (which still seems slow to me), and since 2019 every trust has had access to a real-time data viewer. Knight told the inquiry that MBRRACE has no power to make anyone use it, and that usage varies: some organizations look several times a week, others once every couple of months. An andon cord nobody is required to watch is a decoration.

An andon system is more than the cord

In a plant with a working andon system, an operator who sees something wrong pulls a cord (or, in some plants, pushes a nearby button). Three things have to be true for that to work:

  1. There has to be a way to signal.
  2. The signal has to produce a guaranteed response from someone with authority.
  3. And pulling the cord cannot be professionally dangerous for the person who pulls it.

The operator is not required to diagnose the defect. That's the whole design. If you make people prove they're right before they're allowed to signal, you have built a system where only obvious problems get reported, and obvious problems were never the ones you needed help finding.

The reason you don't require proof is that the costs are lopsided. A false alarm on a production line costs you some downtime. A false accusation of deliberate harm costs a person their reputation, and that is a serious cost worth naming plainly. A missed signal costs patient lives. The imbalance is still wide enough that the person raising the concern should not have to prove the case before protective action starts.

Nobody at the Countess had to conclude Letby was guilty. They only had to move her off the unit while the concern was checked. That is a precaution, not a verdict. The hospital did it in July 2016, after two more babies had died.

Recommendation 9 in the Thirlwall report is essentially that principle written as a safeguarding rule. It calls for a one-page NHS protocol making explicit that “it is irrelevant whether the person to whom the concerns have been expressed does or does not believe they are true, as is the fact that the person raising the concern is not sure,” and that concerns raised in good faith “must be acted upon immediately.”

Thirlwall makes the same point about what already existed in law. Safeguarding legislation, she writes, “requires action on suspicion or concern rather than leaving it to the choice of the person who is concerned. It is not for that person to decide whether or not what they are concerned about has happened, and certainly not for them to be sure it has happened. This was simply not understood at the Countess.”

Consultant pediatricians pulled the cord repeatedly from 2015 onward. Nobody stopped the line.

What happened to the people who raised the concern

Dr. Stephen Brearey, the lead consultant on the neonatal unit, which in US terms makes him the attending physician responsible for it, raised his concerns in a meeting in May 2016. The report says his concerns “had been shouted down by the nurses and ignored by the executives.” The outcome was a decision to wait and see, which the medical director, Ian Harvey, described as “monitor and alert.”

Baby O and Baby P, identified by letter because the babies and their families have court-ordered anonymity, were murdered at the end of June 2016, about six weeks after that meeting.

The important detail comes after the meeting. Brearey emailed his colleagues saying the meeting had been “helpful.” It had not been helpful. He told the inquiry he had wanted it to lead to “something significant in terms of sort of escalation and assurance of safety.” Nothing like that was discussed. The report's assessment:

“The forceful views of the nurses and the absence of any support for his concerns seems to have undermined his confidence in his own views.”

So a consultant with a serious safety concern wrote to his colleagues saying he was satisfied, because the room had made it clear that being unsatisfied was not going to go anywhere.

What followed made it worse. Letby raised a grievance. The outcome, decided by a panel chair from another trust who the report says “read the papers far too quickly” and “came to a view before she began the hearing,” was that the chief executive should apologize to Letby in the presence of her parents, and that Brearey and Dr. Ravi Jayaram were to enter mediation with her and apologize. Failure to do so was to result in disciplinary action. Thirlwall calls the handling of the grievance and its consequences “deplorable.”

Separately, Harvey raised the possibility of a General Medical Council referral with both doctors in an email, framing their cooperation as protection against it. Chief executive Tony Chambers later raised the GMC option in an executives' meeting. As late as March 2017, the executives were still discussing the idea of reporting the two consultants to their regulator.

The hospital's own Speak Out Safely policy said: “no recriminations will follow reports which are made in good faith.” Thirlwall's summary of what staff face across the NHS today, at 4.29, is that many still do not believe they can raise patient safety concerns without career consequences, “including threat of referral to a regulator and threat of disciplinary proceedings. This occurred at the Countess.”

Her plainest version is at 4.40:

“Doctors, with good reason, feared disciplinary action or referral to the GMC.”

I've been working on a book called The Silence Tax, about the price organizations pay when people decide that speaking up is not worth what it might cost them. I'm not going to use this case in it. It's an extreme event and I don't think extreme events teach the everyday lesson well. But the mechanism the report describes is completely ordinary. The consultants weren't silent. Speaking up was made costly for them, and then they were the ones who paid.

Was “witch hunt” a reasonable read at the time?

The accusation was enormous. A nurse with a good reputation, described by her manager as a strong performer, accused by pediatricians of murdering babies, with no forensic evidence and no confession. BBC News, summarizing the report, says the chair of the governance panel had initially described the allegations against Letby as a witch hunt. From the outside, in 2016, I can see how a manager lands there.

I want to resist the hindsight version of this, and so does Thirlwall. She writes at 1.4 that she reminded herself to assess decisions from the position of the person making them at the time, “not what I or they have learnt since.”

That's the right discipline. It is very easy now to look at a decision that turned out catastrophically and call it obviously stupid. We do this constantly in ordinary organizations, and it teaches people that the safest move is to make defensible decisions rather than good ones.

But the defense doesn't survive the standard that already applied. Nobody at the Countess had to believe the doctors. Belief was never the threshold. Suspicion was. Thirlwall's finding on why people froze is one of the sharpest observations in the report:

“Throughout the evidence there was a real sense that people were reticent about saying anything for fear of being wrong. It seems to prevail over the fear of being right.”

On the question of why the director of nursing, Alison Kelly, who was executive lead for safeguarding, did not act, the report answers directly. She “could not believe that one of her nurses would be harming babies.” Harvey “believed there must be a clinical explanation for all the deaths.” Then the line that explains almost everything that followed: “For both of them, the focus was on what else might be causing the deaths and not how can babies be protected.”

Those are two different questions, and they lead to two different sets of actions. The first one you can work on for two years.

The reputation trade did not pay off

There's a notebook entry in the report from Stephen Cross, the hospital's Director of Corporate Affairs, recorded after an inquest into one of the deaths: “Narrative Verdict ‘Unascertained' No negative comments No press.” Thirlwall's thought: “No press meant no damage to the hospital's reputation.”

So how'd that work out for them?

Thirlwall found more broadly that a July 2016 board presentation was

“the beginning of an exercise in spin, steering the Board away from concerns about criminal acts and the necessity of a referral to the police.”

She describes a wider NHS tendency for management to become preoccupied with avoiding blame and protecting reputation.

What the hospital got for all of it is a 1,100-page public inquiry report across three volumes, its former executives named repeatedly, a police investigation into senior leadership, and a lasting association between the hospital's name and the worst thing that happened there.

The reputation calculation is made constantly, in far smaller situations, by people who believe they are protecting their organization. It is almost always the more expensive option, and it almost never looks that way at the time.

Thirlwall says the move to “no blame” was a mistake. I half agree.

Paragraph 4.26 is the one that improvement people should read carefully:

“The move towards a ‘no blame' culture in the NHS, starting in 2000, was a mistake. The focus on system faults rather than the failings of individuals, coupled with a learning culture that omitted acknowledging responsibility or blame, meant difficult conversations about conduct were avoided. This undermined patient safety.”

I agree with more of that than I expected to, and I disagree with the causal story.

To me “no blame” (a short slogan) really means “don't blame individuals for problems caused by the system.” Not even Dr. Deming thought 100% of problems are caused by the system. He estimated the number to be in the low 90s. Which is revolutionary for organizations that assume nearly everything is the fault of an individual.

The part I agree with: “no blame” and “just culture” got collapsed into the same thing, and a lot of organizations now misuse the vocabulary of psychological safety to avoid conversations they should be having. If a person is repeatedly doing something unsafe and nobody will say so, that is not a learning culture. Respect for People includes telling people the truth about their work. I've seen leaders hide behind systems language to avoid an uncomfortable meeting, and I think that's a fair charge. I wish that didn't happen.

The part I don't accept is that systems thinking is what failed at the Countess.

Nothing in the report shows an executive declining to confront someone because they were focused on system faults. Kelly chose not to act because she couldn't believe a nurse would harm babies. Harvey didn't act because he was sure a clinical explanation existed. Chambers told two consultants who had asked him to call the police that they should call the police themselves.

Those are not the errors of people over-committed to system design. Those are the errors of people protecting an institution and avoiding a hard conversation, which is exactly what the report says in the next breath.

There's also a smaller version of the same pattern further down. The report describes Letby as a junior nurse who ignored management instructions she disliked and shouted at a shift leader, with no consequence. An organization that will not have a difficult conversation with a clinician about rule-breaking is not going to have one with its medical director. The failure to confront ran in both directions. Calling that a consequence of no-blame culture gets the direction of causation backwards, at least on the evidence in this report.

Seventeen recommendations, and what they reach

Thirlwall made 17 recommendations. Nearly all of them are structural: in-cot cameras, biometric or CCTV-controlled insulin storage, mandatory lab protocols, national bereavement pathways, interoperable IT by December 2028, board-level monitoring of every child and baby death, a named data lead reporting at least every six months, revised SUDIC guidance, a standing panel of independent experts, more pediatric and perinatal pathologists in training, a barring system for managers with an individual duty of candor, changes to how CQC inspects, and a transfer of the National Guardian's whistleblowing functions to the Parliamentary and Health Service Ombudsman.

Whew. Quite a list.

They fall into two groups. Some make the specific methods of harm harder to repeat: cameras, controlled insulin storage, mandatory lab protocols. Others go at the organizational failure that let concerns go unanswered: escalation routes, data reporting, the safeguarding protocol, manager accountability, how CQC inspects, and where whistleblowing oversight sits.

Two of them look most like an andon system. Recommendation 9 is standard work for a rare, high-consequence concern, and its value is that it removes the judgment call everyone at the Countess got wrong. Recommendation 6 requires a predetermined escalation route for concerning data trends.

Locking the insulin fridge is reasonable, and the same recommendation makes the lab protocol for insulin testing mandatory, which is the countermeasure that would have mattered in August 2015. On its own, none of it does anything about a medical director who is sure there must be a clinical explanation, or a director of nursing who cannot believe one of her nurses would harm a baby.

There is one more tension in the report worth naming. It finds that Jayaram should have reported what he saw in February 2016, and it documents in detail why he and his colleagues didn't:

  • fear of being wrong,
  • fear of the GMC,
  • a grievance process that put them on trial, and
  • an executive team that had already made its position clear.

Both of those things are in the same document. Thirlwall's answer, that suspicion is enough and nobody had to be sure, is correct as a standard. Whether it's a fair expectation of individuals inside a system that was punishing exactly that behavior is a harder question, and I don't think the report fully resolves it.

Recommendations have a poor track record here

Recommendation 17 asks the National Audit Office to audit whether inquiry recommendations get implemented. Early in its work, the inquiry reviewed three decades of recommendations from NHS inquiries and found that most had not been implemented. Thirlwall lists the reasons: no accessible way to track or enforce progress, made worse by constant NHS restructuring, too many recommendations in the first place, unclear ownership, and a lack of political will.

She gives one example that ought to be uncomfortable for anyone who writes corrective actions for a living. In July 2003, Dame Janet Smith's third Shipman report recommended a system of medical examiners so that a doctor's account of a patient's death would face an independent cross-check. The proposal, in Thirlwall's words, “stood still.” Sir Robert Francis supported it again in 2013. It became statutory on 9 September 2024.

That is the countermeasure recommended twelve years before these deaths and delivered eight years after them, and it is the one that might have caught this.

Tracking is also the easy half of the problem. Every trust will be able to show that the cameras are installed and the protocol is on the intranet. Whether any of that changed what happens in a meeting where a consultant says something uncomfortable is much harder to establish, and it is the thing worth knowing.

Anyone who has run a corrective action process knows the difference between confirming an action was completed and confirming it worked. If the audit ends up counting compliance, we will get a report in 2031 saying the recommendations were implemented, and we still will not know whether they made any difference.

So the last recommendation is a check on whether the other sixteen are implemented. There is some irony in it: the report names the sheer volume of recommendations as one reason earlier ones went nowhere, and then adds… checks notes… seventeen more.

Has anyone in leadership been held to account?

As far as I can establish, no Countess of Chester executive or manager has been criminally charged. Cheshire Police arrested three former senior leaders in mid-2025 on suspicion of gross negligence manslaughter, and a further arrest for perverting the course of justice was reported in April 2026. No CPS charging decision has been announced. Alison Kelly was referred to the Nursing and Midwifery Council in 2020 and that case was held pending the police investigation.

NHS managers still cannot be barred from the profession. Doctors and nurses can be struck off by their regulators and blocked from practicing anywhere, but there is no equivalent register for managers, so an executive found unfit at one trust can be hired by another. Thirlwall recommends closing that gap, and it has not been closed.

The government's proposal covers board-level directors and their direct reports, and Thirlwall says the limitation is obvious: it isn't clear how regulating only the most senior managers professionalizes NHS leadership generally. The inquiry, in her words, “heard ample evidence of the ‘revolving door' employment provided for NHS managers about whom serious concerns had been raised.”

The investigation itself is worth pausing on for American readers. England has had a corporate manslaughter law since 2007, which allows prosecution of an organization when the way senior management ran it caused a death. There is no real US equivalent for hospital deaths. What we have instead is the case of RaDonda Vaught, the Vanderbilt nurse convicted in 2022 of criminally negligent homicide after a medication error. No Vanderbilt executive was charged, and the hospital's own failure to report the death was not what the prosecution was about.

I'm not arguing that prosecuting executives is the right countermeasure. I doubt it would change much or save many lives, and Thirlwall's own view is that regulation and barring matter more than criminal cases. But it does say something that the English system can ask whether management decisions were criminal, and ours mostly can't. Or doesn't want to, when you can blame and fire (or convict) an individual.

Every hospital has some version of this. A risk list, a quality dashboard, a board packet. Most have a policy saying concerns raised in good faith will be handled without recriminations. What I'd want to know about any of them is not what the policy says. It's what happened the last time somebody senior was told something that “couldn't be true.”

Get New Posts Sent To You

Select list(s):
Mark Graban
Mark Graban

Mark Graban is an internationally-recognized consultant, author, and professional speaker, and podcaster with experience in healthcare, manufacturing, and startups.

Mark's latest book is The Mistakes That Make Us: Cultivating a Culture of Learning and Innovation, a recipient of the Shingo Publication Award.

He is also the author of Measures of Success: React Less, Lead Better, Improve More, Lean Hospitals and Healthcare Kaizen, and the anthology Practicing Lean.

Mark is also a Senior Advisor to the technology company KaiNexus.

Articles: 5916