4 min read
Designing API integrations that survive the other system
A resilient integration assumes the other system will be slow, will return errors that are not errors, and will occasionally process the same request twice. The three controls that matter most are idempotency keys on every state-changing call, bounded retries with exponential backoff and jitter, and correlation IDs that let you reconstruct what happened afterwards.
- Integration
- Architecture
- Azure
Written by the Codentures engineering team.
Integration work has an unusual property: the most important failure modes are on the far side of a boundary you do not control and cannot fix. The other system will be slow at month end. It will return HTTP 200 with an error in the body. It will rate-limit without documenting the limit. It will be down for an upgrade nobody told you about.
You cannot prevent any of that. You can decide in advance what your system does when it happens, and that decision is most of what separates an integration that runs for years from one that needs a person watching it.
Idempotency, before anything else
Any state-changing call that can be retried must be safe to retry. This is not a nicety: the moment you add retries, and you will, every operation becomes an operation that may be delivered more than once. A network timeout tells you nothing about whether the far side processed the request.
The practical mechanism is an idempotency key generated by the caller, stored by the receiver alongside the result, and returned rather than reprocessed on a repeat. Where you own both sides, implement it properly. Where you do not, find out what the other system's deduplication behaviour is before you rely on it, and if it has none, do the deduplication yourself against a natural key.
// Store the outcome against the key, not just the fact of having seen it;
// a retry needs the original response, not a 409.
var existing = await _store.FindAsync(request.IdempotencyKey, ct);
if (existing is not null)
{
return existing.Response;
}
var response = await _processor.HandleAsync(request, ct);
await _store.SaveAsync(request.IdempotencyKey, response, ct);
return response;Retries that help rather than amplify
Naive retries turn a partial outage into a full one. When a dependency slows down, every caller retries, the dependency slows further, and the retries become the load. Three rules keep retries useful:
- Bound them. A fixed maximum attempt count, and a total time budget shorter than the caller's own timeout. Unbounded retry is a queue with extra steps.
- Back off exponentially, with jitter. Without jitter, everything that failed together retries together, which reproduces the original spike precisely.
- Retry only what is retryable. A 429 or a 503 is worth retrying. A 400 will be a 400 every time; retrying it burns budget and hides the bug.
Above retries sits the circuit breaker: after a threshold of consecutive failures, stop calling for a period and fail fast. It protects the dependency from you and protects your latency from it. In .NET this is a few lines with a resilience pipeline; the engineering is in choosing thresholds you have justified rather than copied.
Move to asynchronous when the coupling hurts
A synchronous call chain means your availability is the product of everyone's availability. Where the business process does not actually require an immediate answer, putting a queue between the systems changes the failure mode from 'the user sees an error' to 'the message waits'.
On Azure this is typically Service Bus for commands that must be processed exactly once in order, Event Hubs for high-volume telemetry-shaped streams, and Azure Functions for the handlers. The important design decisions are not which service to use but what happens to a message that cannot be processed: a dead-letter queue with an actual human process attached to it, or a silent accumulation nobody notices for six weeks.
Validate at the boundary, and be strict about it
Treat every inbound payload as hostile, including from a partner you trust. Most malformed data is a deployment mistake, not an attack, and both need the same handling. Validate against an explicit schema at the edge, reject with a specific error, and never let unvalidated data reach domain logic. Where the payload feeds a downstream system, re-encode for that context rather than passing strings through.
Observability is the deliverable
When an integration misbehaves, the question is always the same: what did we send, what came back, and when? If you cannot answer it in minutes, every incident becomes an archaeology project.
- A correlation ID generated at the edge and propagated through every hop, including into the message payload for asynchronous paths.
- Structured logs with the correlation ID, the operation, the outcome and the duration, with credentials, tokens and personal data redacted at the logging boundary rather than by convention.
- Metrics on the things you would page someone about: error rate, latency percentiles, dead-letter depth, retry volume.
- An alert on dead-letter depth specifically. It is the single highest-signal integration alert there is, and the one most often missing.
None of this is novel. It is simply the set of decisions that has to be made deliberately at design time, because retrofitting idempotency into a system that has been running without it means reconciling the duplicates it has already created.