At around twenty to one on the afternoon of 29 June 2026, the power failed at the data centre that houses a large share of the national digital systems that Digital Health and Care Wales (DHCW) runs for the Welsh NHS. That much is ordinary; power fails. What happened next is not supposed to be possible. The generators did not carry the load. The uninterruptible power supplies — the batteries whose entire purpose is to bridge the seconds between a mains failure and a generator starting — did not bridge anything. A month later, reporting it to the board, the chief operating officer put it in six words:

"None of that worked, and the facility went offline."

Some services had cross-site resilience and rode it out. A small number did not, and those went fully dark. It took until 14 July — the best part of a fortnight — before the organisation declared everything restored. It was the first time in twelve months that DHCW had missed its 99.9 per cent availability target, and it missed it because a building full of redundancy had, when tested, no redundancy at all.

It helps, here, to know what sits behind these systems. DHCW does not run the IT of a single hospital; it runs the national digital infrastructure of the whole of NHS Wales — the systems that every one of the country's GP practices, emergency departments, laboratories and screening services reach into, for a population of some three million people. On the waiting lists alone, an estimated 547,500 patients are waiting on roughly 698,400 open treatment pathways — more pathways than patients, because so many are waiting for more than one thing at once. Each of those pathways is a record, a result, a referral moving through the very systems this data centre keeps alive. When it goes dark, that is not a back-office inconvenience: it is the record a clinician cannot open, the emergency department whose patient system will not admit, the laboratory result that cannot be filed — not in one building but across a nation, at once. The architecture of resilience exists precisely so that a population that size is never exposed to a single point of failure. On 29 June 2026, it was.

There is a detail in the same board meeting that turns this from an accident into something worse. The facility, the board was told, "had two heat related failures" before this one. This was not a bolt from the blue striking a well-run system. It was the third strike on a system that had already told its owners, twice, exactly where it would break.

And there is one further detail, which is where this article begins in earnest. Twelve months before the blackout, DHCW's chief executive had stood in front of the same board, described a near-identical failure, and reached for the most serious term in the health service's vocabulary to describe it. She called it a "never event."

She was right that it should have been. She was wrong that it would be.

What a "never event" is, and why the word matters

In the NHS a Never Event is a term of art, not a figure of speech. It denotes a category of incident so serious, and so completely preventable, that with the correct safeguards in place it simply cannot happen: operating on the wrong part of a patient's body, leaving a surgical instrument inside them, connecting a feeding line to the wrong port. The defining feature of a Never Event is not that it is rare. It is that the barriers which would stop it are known, established, and — where the incident occurs — were not working.

When Helen Thomas borrowed that phrase for a data centre failure, she was making a precise and damning claim about her own organisation's infrastructure. She was saying: this is the sort of failure that a properly commissioned, properly maintained resilient facility makes impossible. If it happened, the safeguards were not there. Thomas is DHCW's founding chief executive; she was the Director of Information at NWIS — the NHS Wales Informatics Service — from 2017, then its interim lead, and carried the organisation across the threshold as NWIS became DHCW in April 2021. The infrastructure she was describing is, in a real sense, the one she has led for the better part of a decade.

Digital Health and Care Wales is the special health authority stood up in April 2021 to carry forward — and, on the evidence, to inherit the unfinished business of — the former NHS Wales Informatics Service. Since March 2025 it has been under formal escalation by Welsh Government — first to Level 3 (Enhanced Monitoring), and, since spring 2026, to Level 4 (Targeted Intervention), the tier immediately below the special-measures regime that governs Betsi Cadwaladr University Health Board. The escalation is about DHCW's ability to deliver. This article is about something adjacent and, in the end, more troubling: its ability to learn.

Because the striking thing about the never event is not that it was said. It is what happened to the words afterwards — and how faithfully the failure they described kept its appointment, year after year, while the organisation reassured its board, edited its record, and pointed at its supplier.

The chain

Set out in sequence, the data centre story is not a run of bad luck. It is a single failure recurring on an almost annual cadence, and at each turn the organisation does the same three things: it reassures the board that the problem is understood and contained; it minimises the incident when it recurs, and edits the sharpest detail out of the published record; and it externalises the cause, locating the fault with the supplier rather than the design. Reassure, minimise, externalise — then recur, worse. Watch it run.

2021. On 11 November a power outage in the data centre disrupts GP practice systems across Wales. It is restored the same day. A warning, logged and survived.

2022. The board is told the move to a second data centre is "probably the final step in that data centre journey" — the step that will finally settle the resilience question and clear the path to the cloud. The phrase belongs to Carwyn Lloyd-Jones: then DHCW's Director of ICT, today its Chief Cloud Officer, and throughout the arc that follows the executive accountable for the data centre estate at the centre of this story. It is his assurance that the journey is, in effect, nearly over. Keep that word — final — in mind; the failures that follow are, each time, presented as the last one.

November 2023. The chief operating officer, Sam Lloyd, is asked about resilience. He tells the board that for the most critical clinical systems, failover between data centres is effectively instantaneous: "clinical critical are our most critical services … generally we're able to instantaneously, or certainly within a few seconds, fail over between data centres, and we've actually demonstrated that as part of the preparation for the data centre move." Within seconds. Demonstrated. In the same period a botched network change during the migration causes an outage across multiple systems and has to be rolled back; in the published minutes, the words "went wrong" and "network outage" do not survive — the cause is reframed as routine documentation and testing delays.

February 2024. DHCW achieves, for the first time, a successful failover of all information services between the two data centres and back. "Geographical resilience," the committee is told. The problem is declared solved.

July 2024. The problem is not solved. A fire-detection system "incorrectly identified a fire" and switched the cooling to a backup system; when staff realised there was no fire and tried to fail the cooling back to the primary, the failback did not work — "a switch that tripped out during the failback," the board was told, caused the outage. The failover — the one demonstrated, the one that worked within seconds — did not hold. Thirty-two services are affected. Total restoration time: five hours and fifty-five minutes. Three services breach their recovery-time SLAs. To the board it is presented as an amber dip in a performance indicator. In her CEO overview at that same meeting, Thomas praises the data centre migration as a "smooth transition" and thanks staff for the "huge undertaking" — without mentioning that the resilience it was supposed to deliver had just failed its first real test. The incident, where it is discussed, comes with a destination for the blame already attached: there are, the board hears, "some questions for them to answer around their maintenance regime." Them being the supplier.

There is, buried in the same July 2024 discussion, one of the few genuine moments of organisational learning in the entire sequence — and it is worth pausing on, because of what happens to it. Lloyd concedes that after the failed network changes the previous November, DHCW had brought in third-party expertise to help, and that "with hindsight we probably should have done that earlier… that's a key lesson for us to take." A key lesson, named as such, on the record.

11 June 2025. The key lesson has not held — and neither has the fix. A false alarm on the fire-detection system, the same trigger as the year before, sets off a chain reaction through the cooling systems; but this time both cooling systems fail at once. The automatic failover — again — does not do what it is supposed to: the mechanism "which should have kicked in effectively didn't." Temperatures spike. Clinical-critical services are recovered by ten that evening by falling over to the second site; full resilience across both data centres is not restored until half past three in the afternoon two days later. Swansea Bay's emergency department continues to feel it in its patient-administration system for days afterwards. The NHS Wales App's development slips because changes have to be postponed to manage the recovery.

And it is here, describing this failure, that Thomas reaches for the phrase:

"Let's not kind of kid ourselves though, this should really be a never event in terms of the level of data centres that we commission."

In the fuller version captured in the recording, she goes on: "there's a lot of work for us to do with the data centre providers to ensure that they can… give us reassurance so this is a never event and it will never happen again."

Now hold that against the published minutes of the meeting. The label — never event — is not in them. Neither is the root cause: the false alarm, the cooling chain-reaction, the failover that should have kicked in and didn't. Neither is the patient impact at Swansea Bay. Neither is the admission, made in the same discussion, that "we did have another incident like this last year" — the July 2024 failure, now acknowledged aloud as a recurrence, and now absent from the record. Thomas's own statement that the "real world impact" of the incident still "needed further translation" — that they did not yet fully understand what it had done to patients — is gone too.

What the board said happened, and what the organisation's permanent record says happened, are two different things. The most serious word anyone used about DHCW's infrastructure that year did not make it into the minutes.

August 2025. The NHS Wales App's first outpatient-appointments pilot, at Hywel Dda, is delivered two weeks late — because key elements were scheduled for the very day of a data centre outage.

November 2025. The board approves extending the data centre contract with CDW Limited from 2026 to 2029, at an additional cost of £1.73 million, to keep the lights on while the cloud migration continues.

29 June 2026. The lights go out. All of them. None of the backups work, and the facility goes offline for a fortnight — the failure Thomas had, a year earlier, said should never happen and would never happen again. For the best part of two weeks, a share of the systems that a nation's clinicians reach for dozens of times a shift was, for some of them, simply not there; only a pre-emptive decision to fail live services across to the second site — described to the board as "fortuitous" — kept the impact from being worse. Reporting it, the chief operating officer offers the board the by-now-familiar coordinates for the fault: "the vast majority of the issues sit with the supplier." The executive accountable for the estate was the same one who, in 2022, had assured the board this was "the final step."

Five years. Five variations on one failure. At every step the organisation had the information it needed — a warning survived, a demonstration that later failed, a key lesson named, a never event declared — and at every step the information failed to change the outcome. The system that was supposed to be resilient was not. And the system that was supposed to make an organisation resilient — its memory, its record, its capacity to carry a lesson from one year into the next — was not either.

And the assurances had names on them. It was Carwyn Lloyd-Jones who told the board in 2022 that the second data centre was the final step. It was Sam Lloyd — the chief operating officer, an executive director of DHCW, and Lloyd-Jones's own line manager — who assured the board in 2023 that clinical systems failed over within seconds and that this had been demonstrated, and who, when they repeatedly did not, located the fault with the supplier in 2024 and again in 2026. And it was Thomas who, in the very summer the failover first failed, told the board the migration had been smooth. The point is not, on this evidence, that anyone lied — an assurance can be sincere and still be overtaken by the facts. It is that at every level of the hierarchy, from the officer who owns the estate to his director to the chief executive, reassurance kept being offered in place of a fix; and that each time the reassurance failed, the record was tidied and the supplier was blamed. Accountability for a resilience that was promised four times and delivered none did not, on the public record, come to rest anywhere.

It is not just the data centre

If the data centre stood alone, it would be a procurement failure and a maintenance failure — serious, but bounded. It does not stand alone. The same pattern — a lesson identified, recorded, and then not retained — recurs across DHCW's portfolio, and the most persuasive witnesses to it are DHCW's own executives, who describe the recurrence out loud, apparently without registering what they are describing.

Consider the NHS number. In February 2026 an independent member, Rowan Gardner, observed that the requirement to use it on all patient correspondence — a patient-safety standard — had been issued "for the fourth time", most recently in a 2015 circular. Issued four times; complied with, evidently, none of them. That circular — 2015/049, "use of the NHS number in all correspondence relating to patient care", described as a safety requirement — had been acknowledged at the committee a year earlier as still not enforced. A safety standard reissued across a decade is not a standard being implemented. It is a lesson the system cannot hold on to.

Or ethnicity data. In April 2025 the director responsible, Ifan Evans, told the audit committee that the standard "has been in the data standard for over a decade, and yet it's not collected in the front-end systems." Over a decade. Still not done.

Or roles and accountability — the very thing a Level 4 escalation exists to fix. A governance review commissioned in 2022 to clarify unclear roles was re-reviewed in 2024 and 2025; and in July 2026, the board's own secretary, Chris Darling, told members that ambiguity of roles and responsibilities "have been a key theme that has come up throughout escalation." Commissioned as a fix in 2022; named as a live problem at the top of the escalation four years later.

Or the funding cycle. In 2022 the chair warned that the annual funding letters were arriving so late as to make planning "extremely difficult." In July 2026, an independent member, David Selway, noted that this was "the second year in a row the remit letter has arrived well after our IMTP had been submitted." It was not a new complaint: a year earlier the director Ifan Evans had already named the late-funding cycle as a structural feature rather than an accident — "the first half is quite turbulent because very often the funding for the year is quite late being confirmed to us." A recurring annual pattern, described as such, by the people living inside it.

The organisation even audits its own inability to learn, and finds it. In March 2026 DHCW's incident-review governance group reported that communication and stakeholder engagement had emerged as "a common theme" not just in one review but "a theme within the… program reports" generally — turning up in the thematic review, the Bridgend disaggregation, and the digital maternity programme alike. Competency and skills gaps surfaced across the same set of reviews. DHCW's learning mechanism, in other words, keeps functioning well enough to re-discover the same unlearned lessons — and then not to learn them.

None of this is CareNHS's characterisation. It is what Gardner said, what Evans said, what Darling said, what Selway said, at meetings held in public. The recurrence is not an inference. It is a matter of record — where the record survives.

Why the loop is broken

An organisation learns from failure the way a person does: it notices what went wrong, it holds the memory, and it lets the memory change what it does next. DHCW's problem can be located precisely along that chain. At the first step it edits what it noticed. At the second it substitutes reassurance for memory. And at the third — where a lesson should change behaviour — the change does not come. Each of these is visible in the record.

The record is edited. The clearest single example is the treatment of the never event, above: a term used aloud, a cause described, a patient impact acknowledged — none of it surviving into the minutes. But it is not a one-off. Comparing DHCW's board recordings against its own published minutes reveals a consistent pattern in which the sharpest material does not make the transition. At the board of 27 November 2025, a triple-source comparison identified eleven separate instances where content present in the meeting was absent from the published minutes — among them the name of a supplier, Atos; Welsh Government's refusal of capital for e-referrals and the integration hub; a candid admission that there would be no "gigantic U-turn" in a twelve-month period; the concession that DHCW "had not yet worked out" what metrics it would measure success by; and a discussion that the escalation criteria themselves "might have to change." Earlier meetings show the same editing at work: a staff-burnout increase of 3.9 per cent, stripped from the July 2025 minutes; the "went wrong" of November 2023, stripped from January 2024's.

Be exact about what this is: a documented difference between what was said and what was published — content present in the recording, absent from the minutes. The detail that would force a lesson is the detail that does not survive. An organisation cannot learn from a failure it has edited out of its own memory.

Assurance is performed. Alongside the editing runs a second mechanism, subtler and in some ways more corrosive: the steady issue of internal-audit reassurance on processes that keep failing. The declarations-of-interest process received "substantial assurance" from internal audit even as that same process failed, across more than a year, to flag the conflicts it existed to catch — a case in which, as the board's own record puts it, "the assurance process itself was deficient." The substantial assurance stamp recurs across cloud, information governance, financial sustainability, the Microsoft agreement. It is worth knowing that this is not a fixed feature: DHCW's own auditors noted that the proportion of reports rated substantial fell from 36 per cent in 2022-23 to 20 per cent in 2024-25, and that risk management itself was downgraded from substantial to reasonable assurance. The reassurance, in other words, is thinning — but the reflex to issue it persists.

Problems are declared solved by deletion. Risks, once mitigated on paper, are removed from the corporate risk register — My Health Online in 2021, two financial-sustainability risks in 2024, four risks in a single quarter in late 2025 — and a risk removed from the register is a problem the board stops watching. Sometimes the timing is its own commentary: the switching service's resilience, celebrated as solved when geographical failover was achieved in early 2024, is the same class of capability that failed that July.

Oversight ends where the learning should begin. In August 2026 an independent member put the structural problem plainly: the programme-delivery committee, he said, "lose oversight when something gets closed out. And… a lot of the value comes after the program's been implemented." Programmes are signed off at closure — the point after which whether they actually delivered any benefit becomes, formally, nobody's job to check. Audit Wales found exactly this with DHCW's flagship internal-transformation programme, Building Our Future, which "closed before long-term benefits were shown." And where the deepest learning ought to be guaranteed — clinical safety — the audit of the Welsh Intensive Care Information System found that at the time of the review there was "no named clinical safety lead on the program board" at all.

And the mechanism built to force the learning failed too. This is the part that should trouble anyone still inclined to read the failures above as bad luck. DHCW's delivery problems were not left to the organisation to fix in the dark. In March 2025 Welsh Government escalated it to Level 3 precisely so that external oversight would drive the improvement the organisation had not managed itself. The escalation is the system's own learning mechanism — the barrier of last resort. It did not correct the problem. DHCW was escalated further, to Level 4. And by July 2026 the board was being told that the improvement plan it had submitted was, in Welsh Government's own words, "not currently supportable." When even the intervention designed to force an organisation to learn ends by escalating rather than correcting, the failure being described is no longer a failure of a data centre, or a programme, or a process. It is a failure of the capacity to learn itself.

The view from outside the room

Everything above is drawn from DHCW's own record — its meetings, its minutes, its audits. The obvious objection to a case built that way is that it is a campaign's reading of the evidence, assembled to a conclusion. So it is worth ending with the one witness who owes the campaign nothing and reached the same place independently.

On 3 August 2026 the Auditor General published Digital Health and Care Wales: Digital Transformation. Its findings, in the auditor's own measured language, describe the same organisation this article has described. DHCW, it found, "does not have its own digital transformation strategy." Its board "receives regular reports, but these mostly cover national programmes", so that oversight of its own internal transformation is weaker and the board cannot closely track progress, costs and benefits. Its risk reports, the auditor noted, "focus on activity rather than showing if those actions actually reduce risk" — a near-perfect description of an organisation that records motion and mistakes it for learning. It "lacks a clear view of the main gaps in its internal digital infrastructure and a clear plan to fix them."

That external portrait is worth holding against the self-portrait DHCW had published only weeks earlier: an annual report claiming 94 per cent of its plan delivered and "no significant control or governance issues". The auditor's picture is not that picture.

The auditor's tally of the national programmes is its own indictment: a laboratory system delayed and over its original 2018 estimate of more than £42 million; a radiology system procurement behind schedule against a programme that ran from 2019 to 2025 at an estimated £60 million-plus; the NHS Wales App delayed; the intensive-care system paused. The report closes with six formal recommendations — on measures and milestones, a costed infrastructure plan, funding clarity, post-closure accountability, skills, and value-for-money evaluation — the bland official vocabulary for an organisation being told, from outside, to build the memory it does not have.

There is a small, sharp irony in one of the auditor's factual findings, and it makes a fitting place to stop. In 2025 DHCW assessed its own digital maturity, using a recognised industry framework, and rated itself at Level 4 — the level, in that framework, at which an organisation "manages through agreed processes, controls and governance arrangements." At the very same time, on an entirely different scale, its regulator had independently assessed it at Level 4 as well: Targeted Intervention. Two different frameworks, the same number, opposite meanings — one the organisation's estimate of how well it governs itself, the other the state's judgement that it cannot be left to. The gap between those two Level 4s is the subject of this article.

An organisation that deletes its failures from its record, reassures itself with audits of processes that keep failing, stops watching problems the moment they are declared solved, and treats the closure of a programme as the end of the question — such an organisation is not merely failing to deliver. It has disabled the faculty by which it might notice, and correct, that it is failing. And this is not an abstract governance defect graded by auditors for other auditors to read. The systems that keep failing are the ones the care of three million people runs through: the record at the bedside, the result in the laboratory, the screening invitation that does or does not arrive. When the memory of an institution fails, it is not only the institution that is exposed.

Helen Thomas was right to call the data centre a never event. The tragedy is that in an organisation which cannot hold its own lessons, a never event is simply an event that has not happened for the last time yet.


Digital Health and Care Wales has been placed under Level 4 (Targeted Intervention) by Welsh Government. CareNHS welcomes a response from DHCW to the matters raised in this article. No response has been received to date. If a response is received, we will publish it in full.

Last reviewed: