Skip to content

Free preview

8.2 · The hazard log

Section 8 · Module 8 · Clinical safety · 17 min

What a hazard actually is

A hazard is a potential source of harm to a patient. Not to a system, not to a service level agreement, not to a project.

That sentence does most of the work in this lesson, because the commonest fault in a hazard log is that it is a risk register wearing a different name.

> "The integration engine fails." — not a hazard. It is a cause. > > "A clinician acts on a result that is two days old, because the feed stopped overnight and nothing > reported it." — a hazard. There is a patient in it, and something happens to them.

The discipline is to keep asking *and then what?* until a patient appears. The engine fails, so results stop, so a ward sees an empty list, so a clinician assumes nothing has been reported, so a critical potassium is not acted on. The last clause is the hazard. Everything before it is the causal chain, which the log also records — but under *cause*, not under *hazard*.

The columns

A hazard log entry is a small structured argument. The columns vary between organisations; the substance does not.

ColumnHolds
IDA stable reference. It will be cited in an incident report years later
HazardThe harm to a patient, in clinical terms
CauseWhat leads to it — usually several, listed separately
EffectWhat the patient experiences
Existing controlsWhat already reduces it, before this project
Initial riskSeverity × likelihood, before new controls
New controlsWhat this project adds — design, procedure, training, monitoring
Residual riskSeverity × likelihood, after
OwnerA named person, not a team
StatusOpen, controlled, closed — and closed needs evidence

Scoring, and the one honest thing about it

Severity and likelihood are each graded on a five-point scale — from *minor* to *catastrophic*, and from *very low* to *very high* — and combined in a matrix to give a risk level of 1 to 5, where the top of the range is unacceptable and the bottom is acceptable.

Two things worth saying plainly about the numbers.

They are a conversation, not a measurement. Nobody knows the true likelihood that a clinician misses a stale result. The value of the matrix is that it forces a clinician and an engineer to disagree explicitly and resolve it, and the record of that disagreement is the real output.

Severity is judged by a clinician. An engineer estimating how bad a duplicate patient record is will usually be wrong, in both directions.

Controls, in order of strength

Not all controls are equal, and a log that treats them as equal is over-claiming.

Design — make the failure impossible. A conditional update that cannot create a patient. Detection — make the failure visible fast. An alert when a feed goes quiet for longer than usual. Procedure — a documented human step. A daily reconciliation. Training — a person knowing to be careful.

Design beats detection beats procedure beats training, every time. A hazard reduced from unacceptable to acceptable purely by training has usually not moved as far as the log says.

Three rows from the St Aidan's ADT hazard log

The interface from module 3, as a safety case would see it.

HAZ-04 · A patient remains recorded as an inpatient after their admission is cancelled

*Cause:* the feed does not carry A11, so the withdrawal of an admission is never sent. *Effect:* the bed appears occupied; the patient appears admitted in systems used by other teams; correspondence and tasks are generated against an episode that did not happen. *Existing controls:* ward clerks correct the bed state manually when they notice.

*Initial risk:* severity considerable, likelihood high — this happens weekly. *New controls:* carry A11 and A13, and route them to every destination that received the A01 (design); a weekly reconciliation of open episodes against the PAS (detection). *Residual:* severity unchanged, likelihood low. *Owner:* named integration lead. *Status:* controlled.

HAZ-09 · A clinician acts on a corrected result believing it is the original

*Cause:* OBX-11 of C is carried but not displayed; the receiving system overwrites silently. *Effect:* a patient may be treated on a value that has been withdrawn, and the record no longer shows what was known when the decision was made. *Existing controls:* none.

*Initial risk:* severity major, likelihood medium. *New controls:* the correction is routed to the same destination as the original and flagged (design); the receiving system displays a correction marker — and this control belongs to the supplier, so the residual rating depends on their delivery. *Residual:* likelihood low, *conditional on that change*, which the log states explicitly. *Status:* open, with a date.

HAZ-11 · A result is filed against the wrong test

*Cause:* a local code is mapped to a nearby but incorrect LOINC concept. *Effect:* the result trends alongside genuine results of a test that was never performed. *Existing controls:* mapping reviewed at build.

*Initial risk:* severity major, likelihood low. *New controls:* unmapped codes are dead-lettered rather than guessed (design); mapping changes require a second reviewer (procedure); a monthly report of codes seen against codes mapped (detection). *Residual:* likelihood very low.

What to notice across the three. Every hazard names a patient consequence. Every one has a control that is design rather than diligence. And HAZ-09 is honest about depending on somebody else — which is the kind of entry that makes a safety case credible, because a log in which every risk is comfortably closed is a log nobody believes.

Is "the integration engine goes down" a hazard?

No. It is a cause, and writing it in the hazard column is the most common way a safety case becomes useless.

A hazard has a patient in it. An outage on its own does not — in fact a loud, total, obvious outage is often among the *safer* failures, because everybody knows something is wrong and falls back to the telephone.

What makes an outage dangerous is what follows, and those are the hazards worth writing:

*A result is not acted on, because a ward sees an empty list and believes nothing has been reported.* *An admission is not visible to the team taking over at handover.* *Messages are delivered hours later, out of order, and a superseded result overwrites a current one.*

Three hazards, one cause, and different controls. The first wants an agreed way for a ward to know the feed is down. The second wants a fallback in the handover process. The third wants ordering guarantees or a replay procedure that respects them. Had "the engine goes down" been the entry, the control would have been "improve engine resilience", which addresses the cause and none of the three.

The technique, which is worth practising deliberately: write the cause, then ask "and then what?" until a person is affected. The row you can stop at is the hazard. Everything above it belongs in the cause column, where it is useful — a hazard with its causes enumerated is exactly how you find that one cause produces three different harms.

Can a hazard be closed by writing a procedure?

Sometimes, and it is much weaker than it looks — so the log has to be honest about which kind of control it is claiming.

A procedure is a real control when four things hold: somebody is named to do it, they are trained, it is recorded when done, and somebody notices when it is not. A daily reconciliation with an owner, a checklist and an exception report is a genuine mitigation.

A procedure is not a control when it is a sentence in a document nobody reads. "Staff will check the patient banner before acting on a result" appears in a great many hazard logs and has, in practice, reduced very little — because it asks people to be reliably vigilant about something rare, which is the one thing humans are worst at.

So the questions to ask of any procedural control: *who does it, how do we know they did, and what happens when they do not?* A control with no answer to the third is training with extra steps.

And the sequencing matters. A procedure should be the control you accept when design and detection have been considered and rejected with reasons. A log that reaches for procedure first is usually one where the engineering was decided before the safety analysis started — which is the wrong order, and the safety case will read like a justification rather than an assessment, because it is one.

The good news is that procedural controls in integration are often replaceable cheaply. A daily manual reconciliation is usually a report somebody could have written in an afternoon, and converting one to the other is the highest-value hour in the project.

What goes wrong

  • Hazards that are causes. No patient in the sentence; controls aimed at the wrong thing.
  • Severity judged by engineers. It is a clinical judgement, which is why the CSO is a clinician.
  • Procedure and training doing all the work. The risk has not moved as far as the log claims.
  • Every row comfortably closed. Nobody believes it, and rightly.
  • No owner, or a team as owner. An unowned control is not a control.

---

Next: lesson 8.3, what happens when an interface harms somebody — the sequence of the first hour, and what your retention policy decided for you months earlier.

That is one lesson of 68

(ECHIA) – EduQan Certified Healthcare Integration Architect runs to 68 lessons across 9 sections, and ends in an assessed, dated certificate you can have verified by anyone.