A request leaves your job runner and nothing comes back. The connection was open, then it was not. Somewhere between your infrastructure and ours a response was written to a socket that had already gone, and you are left holding a question with no answer in it: does that application exist?
Plenty of integrations answer by giving up. The job is marked failed, a line lands in an error channel, and on a good day somebody reads it on Monday. Meanwhile a traveler who was three clicks into buying a visa hears nothing further from anyone. That outcome is expensive, and it comes from a fear that looks entirely reasonable on its face: retry, and you might charge someone twice.
What a timeout actually tells you
It tells you that you stopped waiting. It says nothing about what the server did with the request before you stopped. Networks lose responses about as readily as they lose requests, and the half of the round trip you cannot see is the half that decides whether any work happened.
So the state after a timeout is unknown, and treating an unknown outcome as a failure is a decision. For anything a traveler has already started paying for, it is usually the wrong one.
An integration that never retries converts every transient blip into an abandoned application, and blips are not rare at the scale a booking flow runs at. A deploy on our side or a load balancer recycling a connection is ordinary noise on its own, and either one can strand a traveler who did everything that was asked of them. The volume you lose this way is small enough to be invisible on a dashboard and large enough to be worth fixing.
Retrying is only dangerous when the receiving server cannot tell a retry apart from fresh work, which is a server problem, so we went and solved it on our side.
One key per intent
The applications endpoint honours the standard Idempotency-Key header. The mechanics fit in a paragraph, which is rather the point of them.
Before you send anything, generate a key for the intent you are about to express and write it down. Persisting it first matters more than it sounds: a key created at request time lives only in the memory of the process that is about to die. Store it alongside the record that represents the application in your own database, put it on the request, and send the same one on every retry of that intent.
Two rules then cover everything that can happen.
One extra header on the request: Idempotency-Key: your-unique-string, where the value is anything your system can generate once and reproduce on demand. We keep the stored outcome long enough that any sane retry policy is covered.
The header is the easy half. Deciding what counts as one intent is where integrations go wrong, in two directions. Key on the attempt, with a fresh identifier minted for each retry, and the mechanism does nothing whatsoever: every retry introduces itself as new work and is treated accordingly. Key too broadly, on the traveler for instance, and you quietly block the second application that same traveler legitimately needs next spring.
The scope that works is the one your own database already believes in: one traveler, one trip, one application. If you have a row for it, that row should carry the key.
Dozens of submissions, one application
A partner's job queue redelivered a batch of work after a node came back from the dead. Dozens of identical submissions for the same application arrived at our door inside a few seconds, all of them carrying the same key.
Afterwards there was one application. One charge. The traveler got one confirmation email and never learned that anything unusual had occurred. Nobody on either side opened a ticket, because from the outside nothing did occur. We found it in a graph.
A prevented duplicate charge leaves no trace. The only evidence is a support ticket that never arrives.
Without the key on those requests, the honest description of that afternoon is dozens of live applications against dozens of consular fees, refunded one at a time by somebody who did not plan to spend the day doing it. The government portals would have accepted every one of them. They have no idea your queue hiccupped.
A retry policy worth having
Write the key down before you send. In the same transaction that creates your local record, if your storage allows it. A key held only in memory disappears with the process that was about to crash, and the retry that follows will be indistinguishable from new work.
Back off, and add jitter. Retry on timeouts and on 5xx responses with exponentially growing gaps. Jitter earns its keep when a queue drains after an incident: without it, every worker retries in lockstep and rebuilds the burst that caused the trouble.
Treat a network error after sending as an unknown outcome, not a failure. Your code cannot tell "never arrived" from "arrived, and the reply was lost". Send it again with the same key and let the answer come from us.
Cap the attempts and show a human the last known state. Retrying forever is a way of never admitting that something is wrong. After a handful, stop, and surface the application with whatever you last knew about it so an operator can look instead of guess.
Make your webhook consumers idempotent too. We deliver at least once, so you will occasionally see the same event twice. Dedupe on the event id and every handler becomes safe to replay. Six webhooks we send, three you should actually handle covers which ones are worth wiring up first.
Exactly-once delivery over a network does not exist. What exists is at-least-once delivery arriving at a receiver that recognises repeats, and for every practical purpose you have, the two behave the same. That is the entire trick, and it is enough. The retry stops being a decision anyone has to make, and a redelivered queue turns into a line on a graph nobody has to explain afterwards.