Skip to content

Free preview

9.2 · Retry, backoff and ordering

Section 9 · Module 9 · Delivery · 17 min

Classify first, then decide

Retry is not a setting. It is a response to a classification, and getting the classification wrong produces either lost messages or a retry storm.

Transient — it might work later. A refused connection, a timeout, a 503, a locked file, a database deadlock. Retry.

Permanent — it will never work. A malformed message, a rejected schema, a 400, an unknown account. Do not retry; retrying a message the destination will always refuse simply moves the failure later and fills a queue on the way.

Neither — a credential or permission failure, a 401, an expired certificate. It will not fix itself and it will not be fixed by waiting. Alert, and stop trying.

Most engines default to retrying everything, which is the safest default and the wrong policy. Sorting failures into those three buckets is the actual design work.

Backoff, and the storm

Retry with no delay is a denial of service you built yourself. A destination that is struggling receives every failed message again immediately, plus the new traffic, and the thing that was slow becomes the thing that is down.

Exponential backoff — one second, two, four, eight — gives a struggling system room. A cap stops the intervals becoming absurd. And jitter, a random spread, matters more than it sounds: a hundred messages that failed together and back off identically will retry together, which is the same storm with extra steps.

Then a limit. Retry for ever is not a policy; it is a queue that never drains and an alert nobody receives. Bound it, and decide what happens at the end.

The dead letter

Where a message goes when delivery has permanently failed or the retries are exhausted.

It must be durable, inspectable, and replayable, and it must contain more than the payload: which destination, which attempt, the last error, and the correlation identifier from lesson 8.3. A dead-letter queue that holds only payloads is an archive, not a recovery mechanism.

And it must be watched. A dead letter is the engine saying *a person is needed*, and an unwatched dead-letter queue is exactly lesson 7.3's mailbox that nobody owned.

Retry breaks order

This is the part people discover late.

Message A fails and is retried. Message B, sent afterwards, succeeds immediately. A arrives after B — and if both concern the same entity, the older one has just overwritten the newer.

Three answers, in increasing cost.

Do nothing, because order does not matter. True for most feeds — independent orders, independent tracking events. Establish it rather than assume it.

Preserve order by blocking. A failed message holds the queue until it succeeds or is dead-lettered. Order is guaranteed and one bad message stops the feed, which is a real trade and sometimes the right one.

Order per entity. Messages about the same order are delivered in sequence; messages about different orders proceed independently. The best answer where it is available, and it needs a key — which is what module 7 was for.

And the cheapest defence of all, which lives at the other end: carry a version or a timestamp, and have the destination ignore anything older than what it holds. Then late delivery is harmless, and you have not had to make the pipeline strictly ordered to get a correct result.

The price that went back up

Meridian sends price updates from Atlas to the customer portal. Independent messages, one per product, several thousand a day.

14:02. Product EQ4471 is repriced from £890 to £845 for a promotion. The message is sent. The portal's API is briefly overloaded and returns 503. The channel queues it and backs off.

14:03. A correction: the promotional price should be £849, not £845. Sent, delivered immediately — the overload has passed.

14:07. The first message's fourth retry succeeds. £845 is written over £849.

What the portal shows. £845. Which is wrong, is lower than intended, and stays that way for eleven days until somebody in finance notices the margin on that line.

Why nothing detected it. Both deliveries succeeded. The channel's history shows two successful deliveries in the correct order of *attempts*. There is no error anywhere, and the portal has no reason to think anything is odd — it received a price and stored it.

Three fixes, and they are not equivalent.

*Order per entity.* Messages about the same product are delivered in sequence. Correct, and it costs a keyed queue and some throughput.

*Block the queue on failure.* Guarantees order for everything, and one stuck message stops every price update in the estate. Too strong here.

*A version on the message, checked at the destination.* Atlas stamps each price update with a monotonic version; the portal ignores anything not newer than what it holds. Late delivery becomes harmless, retries stay free to reorder, and no throughput is lost.

The third was chosen, and the reason is worth generalising: it makes the *correct* outcome independent of delivery order, rather than making delivery order correct. Every other approach in this lesson is an attempt to control the pipeline; this one removes the pipeline's ability to cause the problem.

The design question underneath, which should be asked of every feed at build time: *if two messages about the same thing arrive in the wrong order, what happens?* If the answer is "the wrong one wins", you need one of these three before go-live, not after.

Should you retry a message the destination rejected with a 400?

No. It is a permanent failure, and retrying it is a queue you are filling for no reason.

A 400 or 422 means the destination has looked at the message and decided it is wrong. That opinion will not change in thirty seconds. Retrying achieves three unhelpful things: it delays the moment anybody finds out, it consumes retry capacity that transient failures need, and it can look like an attack to a partner watching their own logs.

The correct response is to treat it as terminal: dead-letter it, with the destination's own error text attached, and alert. Then somebody can read the reason — usually a specific field — and decide whether it is your mapping or their data.

Two cautions, because the rule is not quite universal.

Some systems return the wrong code. A 400 that actually means "I am overloaded" exists in the wild, and if a partner does that you will discover it as a pile of dead letters that would have succeeded. Where you know a partner is like this, classify per partner rather than per code, and write down why.

A 409 conflict is usually about state, not about the message. It may well succeed later, and it frequently means the message is a duplicate — which is a success wearing the wrong hat.

And the reverse error is worth naming too: do not treat a 500 as permanent. A destination's internal error is very often transient, and dead-lettering it immediately turns a five-minute blip into a manual replay.

Does your feed need ordered delivery?

Most do not, and the ones that do usually need it only per entity — but the question has to be asked before go-live rather than after.

The test: if two messages about the same thing arrive in the wrong order, is the result wrong?

For independent facts — separate orders, separate tracking events, separate despatches — the answer is no, and strict ordering would cost throughput for nothing.

For updates to the same entity, the answer is yes, and it is the worked example: a stale price overwriting a new one, a status going backwards, an address reverting.

Three ways to get a correct result, and only one of them is about ordering the pipeline.

Per-entity ordering. Keyed queues, so messages about one order are sequential and different orders are parallel. Needs a key and some engine support.

Global ordering by blocking. Simple, strong, and one bad message stops everything. Reserve it for feeds where the whole stream is one sequence — a ledger, a replicated log.

Version checking at the destination. Carry a monotonic version or an authoritative timestamp, and have the receiver discard anything older than what it holds. This makes order irrelevant, which is better than making it correct.

Prefer the third where the destination can be influenced, the first where it cannot, and the second almost never.

And record the answer in the specification. "Ordering is not guaranteed; the destination must apply the version" is a sentence that prevents the incident, and it also tells the destination's team what they are responsible for — which they will otherwise assume is nothing.

What goes wrong

  • Retrying everything. Permanent failures fill the queue and delay the alert.
  • No jitter. A hundred messages that failed together retry together.
  • Unwatched dead letters. The engine asked for a person and nobody came.
  • Assuming order. A stale price overwrites a new one and both deliveries succeeded.

---

Next: lesson 9.3, the senders — and what "success" actually means for each one, which is different every time.

Lab E25 makes a destination fail intermittently and proves a version check beats an ordered queue.

That is one lesson of 94

(ECSIA) – EduQan Certified Systems Integration Engineer runs to 94 lessons across 12 sections, and ends in an assessed, dated certificate you can have verified by anyone.