When repeated requests make an outage worse

Timeouts, bounded retries and idempotency. Prevent a small dependency failure from spreading through the whole system.

An open laptop displaying code in a dim workspace
Photography: Aristo Rinjuang / Unsplash

An external API slows down and your application starts retrying requests. That sounds like a sensible safeguard. But if every client does it at the same time, an overloaded service receives even more work. Retries need rules: which operations qualify, how many attempts are allowed and when to stop for good.

A timeout sets the limit on waiting

Every remote call needs bounded connection and response times. Consider the total duration of the user operation as well. If a customer waits for three services, three independent long timeouts can create an unacceptable delay. The time budget must cover the whole journey, not just one function.

A timeout does not tell you whether the other system completed the operation. The response might have been lost after a successful write. That distinction matters when charging a payment, creating an order or sending a message.

Retry only operations that are safe to repeat

Temporary unavailability may justify another attempt. Invalid input or insufficient permissions usually will not. For writes, use the idempotency mechanism supported by the API. The same key must represent the same intended operation; its meaning must not change between attempts.

Limit attempts and spread them over time

Increase the wait between attempts and add random jitter so clients do not all return at the same instant. Set both an attempt limit and an overall deadline. If the service provides Retry-After, respect it within the available time budget.

Check for retries at multiple layers. If the SDK, backend and queue each repeat a request, the number of calls can multiply. Establish which layer owns retries and measure original requests separately from additional attempts.

Decide how to behave without the dependency

Not every function needs every supplier. A catalogue can remain available without recommendations. For an essential dependency, provide a clear error and a safe next step. A circuit breaker can temporarily stop calls to a repeatedly failing service, but its recovery behaviour needs careful design.

  • Set deadlines for the call and the complete operation.
  • Verify idempotency for writes.
  • Bound retries and introduce jitter.
  • Test slow responses, lost responses and complete failure.

What to take away

Retries do not replace availability. They are bounded assistance for temporary failures, giving a service room to recover without multiplying its workload.

Documentation and further reading

Mgr. Martin Hlavaj, MBA

Software Engineer

All articles

Hear about an outage early.

Add your website or API to UpBot and choose who receives the alert.

Start monitoring for free