
Building a Typed Retry Policy for Rust HTTP Clients
Why a typed retry policy matters
A retry system usually needs to answer three questions:
- Should this request be retried?
- How long should we wait before retrying?
- How many attempts are allowed?
If these rules are encoded as plain booleans and integers everywhere, the code becomes hard to audit. A typed policy gives you a single place to express the rules and makes invalid states harder to represent.
This is especially useful for:
- HTTP clients calling flaky upstream services
- background jobs that can be safely repeated
- rate-limited APIs that return transient failures
- services that need exponential backoff with jitter
The goal is not to build a full networking framework. Instead, we’ll create a compact policy object that can be plugged into reqwest or any other client.
Design goals
A good retry policy should be:
- Explicit: retry conditions are visible in one place
- Safe: non-idempotent requests are not retried accidentally
- Composable: backoff and retry rules can evolve independently
- Testable: the policy can be unit-tested without network calls
We will model the policy with a few small types:
RetryPolicyfor the overall configurationRetryDecisionfor the result of evaluating an error or responseBackofffor delay calculationRetryableMethodfor request safety
Core types
Here is a compact implementation of the policy layer.
use std::time::Duration;
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum RetryDecision {
RetryAfter(Duration),
DoNotRetry,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum RetryableMethod {
Idempotent,
NonIdempotent,
}
#[derive(Debug, Clone)]
pub struct RetryPolicy {
max_attempts: u32,
base_delay: Duration,
max_delay: Duration,
}
impl RetryPolicy {
pub fn new(max_attempts: u32, base_delay: Duration, max_delay: Duration) -> Self {
assert!(max_attempts >= 1);
assert!(base_delay <= max_delay);
Self {
max_attempts,
base_delay,
max_delay,
}
}
pub fn max_attempts(&self) -> u32 {
self.max_attempts
}
pub fn decide(
&self,
attempt: u32,
method: RetryableMethod,
status: Option<u16>,
is_network_error: bool,
) -> RetryDecision {
if attempt >= self.max_attempts {
return RetryDecision::DoNotRetry;
}
if method == RetryableMethod::NonIdempotent {
return RetryDecision::DoNotRetry;
}
let retryable_status = matches!(status, Some(429) | Some(500) | Some(502) | Some(503) | Some(504));
if is_network_error || retryable_status {
let delay = self.backoff_delay(attempt);
RetryDecision::RetryAfter(delay)
} else {
RetryDecision::DoNotRetry
}
}
fn backoff_delay(&self, attempt: u32) -> Duration {
let factor = 1u64 << attempt.min(10);
let millis = self.base_delay.as_millis() as u64 * factor;
let delay = Duration::from_millis(millis);
delay.min(self.max_delay)
}
}This version keeps the policy simple and predictable:
- retries are limited by
max_attempts - only idempotent requests are retried
- common transient HTTP statuses are retried
- network errors are treated as retryable
- backoff grows exponentially and is capped
Why method safety should be typed
Not every HTTP method is safe to retry. GET, HEAD, PUT, and DELETE are usually idempotent, while POST often is not. You can represent that distinction explicitly rather than relying on comments or convention.
For example:
fn classify_method(method: &reqwest::Method) -> RetryableMethod {
match *method {
reqwest::Method::GET
| reqwest::Method::HEAD
| reqwest::Method::PUT
| reqwest::Method::DELETE
| reqwest::Method::OPTIONS => RetryableMethod::Idempotent,
_ => RetryableMethod::NonIdempotent,
}
}This is a conservative default. In real systems, some POST requests are safe to retry if they carry idempotency keys or are otherwise designed for repetition. If that applies to your API, you can extend the enum with a third state such as IdempotentWithToken.
| Method class | Example methods | Default retry stance |
|---|---|---|
| Idempotent | GET, PUT, DELETE | Safe to retry if the error is transient |
| Non-idempotent | POST, PATCH | Do not retry by default |
| Custom safe semantics | POST with idempotency key | Retry only with explicit policy support |
Adding jitter to avoid retry storms
Pure exponential backoff can cause many clients to retry at the same time. Adding jitter spreads retries out and reduces load spikes.
A simple approach is to randomize the delay between 50% and 100% of the computed backoff. The rand crate is a practical choice for this.
use rand::Rng;
use std::time::Duration;
fn with_jitter(delay: Duration) -> Duration {
let mut rng = rand::thread_rng();
let millis = delay.as_millis() as u64;
let lower = millis / 2;
let upper = millis.max(1);
Duration::from_millis(rng.gen_range(lower..=upper))
}You can then apply jitter inside decide:
let delay = with_jitter(self.backoff_delay(attempt));
RetryDecision::RetryAfter(delay)Jitter is especially important when many workers share the same failure mode, such as a database outage or a throttled API.
Integrating the policy with reqwest
The policy becomes useful when paired with an HTTP client. The following example shows a retry loop around a reqwest request.
use reqwest::{Client, Method, StatusCode};
use std::time::Duration;
use tokio::time::sleep;
#[derive(Debug)]
enum RequestError {
Http(reqwest::Error),
RetryLimitExceeded,
}
async fn send_with_retry(
client: &Client,
method: Method,
url: &str,
policy: &RetryPolicy,
) -> Result<String, RequestError> {
let retryable_method = classify_method(&method);
for attempt in 0..policy.max_attempts() {
let response = client
.request(method.clone(), url)
.send()
.await;
match response {
Ok(resp) => {
let status = resp.status();
if status.is_success() {
return resp.text().await.map_err(RequestError::Http);
}
let decision = policy.decide(
attempt,
retryable_method,
Some(status.as_u16()),
false,
);
match decision {
RetryDecision::RetryAfter(delay) => {
sleep(delay).await;
continue;
}
RetryDecision::DoNotRetry => {
return Err(RequestError::RetryLimitExceeded);
}
}
}
Err(err) => {
let decision = policy.decide(attempt, retryable_method, None, true);
match decision {
RetryDecision::RetryAfter(delay) => {
sleep(delay).await;
continue;
}
RetryDecision::DoNotRetry => return Err(RequestError::Http(err)),
}
}
}
}
Err(RequestError::RetryLimitExceeded)
}This loop is intentionally straightforward:
- successful responses are returned immediately
- retryable status codes trigger a delay and another attempt
- network errors are retried only for safe methods
- the loop stops after the configured limit
In a production implementation, you would likely return richer error information, including the final status code or the last network error.
Handling 429 responses correctly
A common mistake is to retry 429 Too Many Requests using only local backoff. Many APIs include a Retry-After header, which should take precedence when present.
You can extend the policy to honor that header:
fn retry_after_from_response(resp: &reqwest::Response) -> Option<Duration> {
let value = resp.headers().get(reqwest::header::RETRY_AFTER)?;
let text = value.to_str().ok()?;
if let Ok(seconds) = text.parse::<u64>() {
return Some(Duration::from_secs(seconds));
}
None
}Then, when a 429 is returned, prefer the server-provided delay:
let delay = retry_after_from_response(&resp)
.unwrap_or_else(|| policy.backoff_delay(attempt));This is a good example of why typed policy logic is useful: the retry loop can remain generic while special-case behavior stays isolated.
Testing the policy
Because the decision logic is pure, it is easy to test without any network access.
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn does_not_retry_non_idempotent_methods() {
let policy = RetryPolicy::new(3, Duration::from_millis(100), Duration::from_secs(2));
let decision = policy.decide(
0,
RetryableMethod::NonIdempotent,
Some(503),
false,
);
assert_eq!(decision, RetryDecision::DoNotRetry);
}
#[test]
fn retries_transient_status_codes() {
let policy = RetryPolicy::new(3, Duration::from_millis(100), Duration::from_secs(2));
let decision = policy.decide(
0,
RetryableMethod::Idempotent,
Some(503),
false,
);
match decision {
RetryDecision::RetryAfter(delay) => {
assert!(delay >= Duration::from_millis(100));
}
RetryDecision::DoNotRetry => panic!("expected retry"),
}
}
#[test]
fn stops_after_max_attempts() {
let policy = RetryPolicy::new(2, Duration::from_millis(100), Duration::from_secs(2));
let decision = policy.decide(
2,
RetryableMethod::Idempotent,
Some(503),
false,
);
assert_eq!(decision, RetryDecision::DoNotRetry);
}
}Good tests for retry logic should cover:
- attempt limits
- idempotent vs non-idempotent behavior
- retryable and non-retryable status codes
- backoff growth and capping
- server-provided delays when applicable
Practical best practices
A retry policy is only useful if it matches the behavior of the upstream system. Keep these rules in mind:
- Retry only transient failures: network timeouts, connection resets, and 5xx responses are typical candidates.
- Avoid retrying validation errors:
400,401,403, and404usually indicate a permanent problem. - Respect idempotency: never retry unsafe operations unless the API explicitly supports it.
- Cap delays: exponential backoff should not grow without bound.
- Add jitter: this prevents synchronized retry bursts.
- Log retry decisions: include attempt number, delay, and reason for observability.
- Expose policy configuration: different services often need different limits.
A useful production pattern is to make the policy configurable through application settings while keeping the decision logic itself fixed and testable.
Extending the model for real systems
Once the basic policy works, you can extend it in several directions:
- add a
RetryReasonenum to distinguish status codes, timeouts, and connection failures - support per-endpoint policies
- allow custom predicates for API-specific retry rules
- integrate cancellation so retries stop when a request is no longer relevant
- emit metrics for retry count, success after retry, and final failure rate
These extensions are easier when the core design is already typed and centralized.
Conclusion
A typed retry policy makes HTTP retry behavior easier to reason about and safer to evolve. By separating retry decisions, method safety, and backoff calculation, you avoid scattered logic and reduce the risk of retrying requests that should not be repeated.
The pattern is small enough to implement quickly, but it scales well as your client code grows. If you need robust request handling in Rust, this is a strong foundation.
