Classify each failure as transient or terminal before you retry, retry by re-enqueueing a Queueable with a delay rather than looping, send an idempotency key on every write, and set the callout timeout from the endpoint's measured p99, not the 120-second maximum.
The endpoint is up. Postman gets a 200. The integration works nine times out of ten. And every few hours an exception email arrives with System.CalloutException: Read timed out, or a 503 nobody can reproduce, or a record that quietly never synced. Intermittent callout failures are the hardest integration problem on the platform because the evidence is spread across three systems and the retry you added last sprint may be hiding the cause.
This article is a diagnostic method and a reference design. It assumes you already know how to make a callout. Every Apex sample states the API version it was written against, and every claim about the platform names its source in the references at the end. The redirect case, where a 301 or 302 arrives instead of a 200, has its own article: Why 301 Redirects Break Salesforce Callouts.
Why do callouts fail intermittently when the endpoint looks healthy?
Because “healthy” is measured from somewhere else. A callout fails when any one of five things crosses a threshold during that specific request. The five are the response time against your timeout, the upstream’s capacity against its rate limit, the network path, the governor limits of the transaction you are in, and the order of operations inside that transaction.
The platform constraints that matter, from the Apex Developer Guide’s governor limits, its Callout Limits and Limitations page and the HttpRequest reference:
- A single callout waits as long as
HttpRequest.setTimeoutallows: 10 seconds by default, 1 millisecond minimum, 120,000 milliseconds maximum. - A transaction may make at most 100 callouts, and all callouts in a transaction share a cumulative timeout of 120 seconds. These limits are the same in synchronous and asynchronous Apex.
- Callouts are not allowed while a transaction has uncommitted work pending. Do your DML after the callout, or in a separate transaction.
- A synchronous request that runs longer than 5 seconds counts against the org’s limit on concurrent long-running requests, which Salesforce now sets at one per 100 applicable licences, with a minimum of 10 and a maximum of 50. Older material quotes a flat 10. The governor limits page excludes HTTP callout processing time from that 5-second count; the Apex around the callout still runs on the clock.
- The endpoint’s host must be covered by a Remote Site Setting, or the callout must use a Named Credential, which needs no Remote Site Setting.
An endpoint with a p99 of 12 seconds and a default 10-second timeout fails about one request in a hundred, every hour, forever, while every dashboard stays green. That is what “intermittent” usually means.
What do the failures look like in the logs?
Each cause leaves a different signature. Match the text in your exception email or debug log to the list below before you change anything.
System.CalloutException: Read timed out
The endpoint did not answer within the timeout you set, or the 10-second default you did not set. Salesforce’s knowledge article on callout timeouts says a timeout surfaces as a System.CalloutException and that setTimeout is where you change the wait. The message text above is as it appears in my projects’ logs.
System.CalloutException: Exceeded maximum time allotted for callout (120000 ms)
In my projects this message has meant the 120,000-millisecond ceiling: either one callout set to the maximum, or several callouts whose combined wait reached the transaction’s 120-second cumulative budget. The wording above is as it appears in my logs; the figures are from the Callout Limits and Limitations page.
System.LimitException: Too many callouts: 101
The 101st callout in one transaction. It is almost always a loop over records, or a Batch Apex scope larger than 100 with one callout per record. The Apex Reference Guide lists LimitException among the exceptions that cannot be caught.
System.CalloutException: You have uncommitted work pending. Please commit or rollback before calling out
DML ran before the callout in the same transaction. Salesforce’s knowledge article on this error is explicit that callouts are not allowed with an uncommitted transaction pending, and that the fix is to restructure so callouts happen first or in another context.
System.CalloutException: Unauthorized endpoint, please check Setup->Security->Remote site settings. endpoint = https://...
The host is not registered. The Metadata API reference for RemoteSiteSetting states the rule: before an Apex callout can call an external site, the site must be registered or the call fails. The wording of the message above is as it appears in my projects’ logs.
HTTP 429, 502, 503, 504 where a 2xx was expected
These are not exceptions. Apex hands you the HttpResponse whatever the status code, and your code decides. If nothing checks getStatusCode(), the failure is silent: the callout “succeeded” and the record was never processed downstream. A 429 is the upstream telling you to slow down, and it may carry a Retry-After header you should honour.
Top Tip: Search your exception emails for the exact strings above. The split between “Read timed out” and “Exceeded maximum time allotted” alone tells you whether to tune one timeout or restructure a transaction.
What are the options when a callout fails, and what do they cost?
You have four moves, and the right one depends on whether the failure is transient. Blind retry, where every failure is retried the same way, is the most common design and the one that causes the most damage.
| Option | When it fits | What it costs |
|---|---|---|
| Retry in the same transaction | Idempotent reads against an endpoint with brief blips | Each retry eats the 120-second cumulative budget and the 100-callout limit; you cannot wait between attempts |
| Re-enqueue through a Queueable with an attempt counter | Writes and anything that needs a real delay | Async plumbing, an attempt counter, and a stable idempotency key so the second write is not a duplicate |
| Park and alert | Terminal failures: 4xx other than 429, bad payloads, authentication | A failure record or log, and someone who reads it |
| Stop and fail fast | Repeated failures to one endpoint during an outage | A circuit-breaker state and a decision about what the user sees meanwhile |
Classified retry means you decide the move from the status code or exception first. Timeouts, 429 and 5xx go to the retry path; everything else parks. A blind retry of a 400 is a bug you run three times.
What architecture should you choose, and why?
Classify first, retry by scheduling rather than looping, and make every write idempotent. The parts:
- A classifier turns a response or exception into one of three outcomes: success, retry, or fail. It is the one place the 429 and 5xx rules live.
- Bounded retry with backoff by re-enqueueing. Apex has no sleep call, so a delay is a scheduled job, not a loop.
System.enqueueJobtakes a delay in minutes from 0 to 10, and Salesforce’s Spring ‘23 developer blog names exactly this use: slowing chained jobs that make rapid callouts to a rate-limited system. A Queueable may enqueue one child job, which is all a retry chain needs.AsyncOptions.MaximumQueueableStackDepthcaps the chain if a bug tries to make it infinite, and overrides the default limit of five in Developer and Trial Edition orgs. In the guide’s sample the first job runs at depth 1, so a cap equal to the attempt count fits. - An idempotency key on writes. The retry must send the same key as the first attempt, so a request that timed out after the server processed it does not create a second order.
Idempotency-Keyis a proposed standard header, an IETF Internet-Draft whose current revision has expired, and it is the header Stripe’s API uses for exactly this purpose. - A timeout budget. Per callout, set from the endpoint’s measured p99 plus a margin, never the 120-second maximum. Per transaction, add them up and keep the total well under 120 seconds.
- A correlation ID header on every request, logged on both sides, so a failure in Salesforce can be matched to one line in the downstream system’s logs.
How do you implement it in Apex?
Four classes: a client that sets the timeout and headers, a classifier, a Queueable that retries by re-enqueueing with a delay, and a test that scripts timeouts, 429 and 503 through an HttpCalloutMock. The samples use a Named Credential called Order_Service, and park failures in a custom object Callout_Failure__c with four text fields: Record_Id__c, Correlation_Id__c, Reason__c and Status_Code__c. Substitute your own logging object or framework.
The classifier:
// Written against API version 67.0 (Summer '26)
public with sharing class CalloutOutcome {
public enum Kind { SUCCESS, RETRY, FAIL }
public final Kind outcomeKind;
public final Integer statusCode;
public final String reason;
public final Integer retryAfterSeconds;
private CalloutOutcome(Kind outcomeKind, Integer statusCode, String reason, Integer retryAfterSeconds) {
this.outcomeKind = outcomeKind;
this.statusCode = statusCode;
this.reason = reason;
this.retryAfterSeconds = retryAfterSeconds;
}
public static CalloutOutcome fromResponse(HttpResponse res) {
Integer code = res.getStatusCode();
if (code >= 200 && code < 300) {
return new CalloutOutcome(Kind.SUCCESS, code, null, null);
}
if (code == 408 || code == 429 || code == 502 || code == 503 || code == 504) {
return new CalloutOutcome(Kind.RETRY, code, 'HTTP ' + code, parseRetryAfter(res.getHeader('Retry-After')));
}
if (code == 301 || code == 302 || code == 307 || code == 308) {
return new CalloutOutcome(Kind.FAIL, code, 'Redirect to ' + res.getHeader('Location'), null);
}
return new CalloutOutcome(Kind.FAIL, code, 'HTTP ' + code, null);
}
public static CalloutOutcome fromException(Exception e) {
String message = e.getMessage();
Boolean isTimeout = e instanceof System.CalloutException
&& message != null
&& (message.contains('Read timed out') || message.contains('Exceeded maximum time allotted'));
if (isTimeout) {
return new CalloutOutcome(Kind.RETRY, null, message, null);
}
return new CalloutOutcome(Kind.FAIL, null, e.getTypeName() + ': ' + message, null);
}
private static Integer parseRetryAfter(String header) {
if (String.isBlank(header) || !header.isNumeric()) {
return null;
}
return Integer.valueOf(header);
}
}
The client, with the timeout set below the endpoint’s measured p99 and both headers on every request:
// Written against API version 67.0 (Summer '26)
public with sharing class OrderSyncClient {
// Set from the endpoint's measured p99 plus margin. The platform maximum is 120,000 ms.
private static final Integer TIMEOUT_MS = 20000;
public static CalloutOutcome push(String payload, String idempotencyKey, String correlationId) {
HttpRequest req = new HttpRequest();
req.setEndpoint('callout:Order_Service/v1/orders');
req.setMethod('POST');
req.setTimeout(TIMEOUT_MS);
req.setHeader('Content-Type', 'application/json');
req.setHeader('Idempotency-Key', idempotencyKey);
req.setHeader('X-Correlation-Id', correlationId);
req.setBody(payload);
try {
HttpResponse res = new Http().send(req);
return CalloutOutcome.fromResponse(res);
} catch (System.CalloutException e) {
return CalloutOutcome.fromException(e);
}
}
}
The Queueable. The callout happens before any DML, the idempotency key is created once and carried through every attempt, and the backoff honours Retry-After when the upstream sends it. The chain is cut short under test so that each test checks one attempt and its scheduled delay. The Apex Developer Guide says chained jobs can be tested with an appropriate stack depth and that the delay is ignored during Apex testing, which is why the test asserts the recorded delay rather than timing anything.
// Written against API version 67.0 (Summer '26)
public with sharing class OrderSyncJob implements Queueable, Database.AllowsCallouts {
private static final Integer MAX_ATTEMPTS = 4;
private static final Integer MAX_DELAY_MINUTES = 10; // the ceiling System.enqueueJob accepts
@TestVisible private static Integer lastScheduledDelayMinutes;
private final Id recordId;
private final Integer attempt;
private final String idempotencyKey;
public OrderSyncJob(Id recordId) {
this(recordId, 1, UUID.randomUUID().toString());
}
private OrderSyncJob(Id recordId, Integer attempt, String idempotencyKey) {
this.recordId = recordId;
this.attempt = attempt;
this.idempotencyKey = idempotencyKey;
}
public void execute(QueueableContext context) {
String correlationId = idempotencyKey + '/' + attempt;
String payload = '{"recordId":"' + recordId + '"}';
CalloutOutcome outcome = OrderSyncClient.push(payload, idempotencyKey, correlationId);
if (outcome.outcomeKind == CalloutOutcome.Kind.SUCCESS) {
return;
}
if (outcome.outcomeKind == CalloutOutcome.Kind.RETRY && attempt < MAX_ATTEMPTS) {
Integer delay = Math.min(MAX_DELAY_MINUTES, backoffMinutes(attempt, outcome.retryAfterSeconds));
lastScheduledDelayMinutes = delay;
if (!Test.isRunningTest()) {
System.enqueueJob(new OrderSyncJob(recordId, attempt + 1, idempotencyKey), delay);
}
return;
}
insert new Callout_Failure__c(
Record_Id__c = recordId,
Correlation_Id__c = correlationId,
Reason__c = outcome.reason,
Status_Code__c = String.valueOf(outcome.statusCode)
);
}
private static Integer backoffMinutes(Integer attempt, Integer retryAfterSeconds) {
if (retryAfterSeconds != null) {
return Math.max(1, Math.ceil(retryAfterSeconds / 60.0).intValue());
}
return Math.pow(2, attempt - 1).intValue(); // 1, 2, 4 minutes
}
}
Starting the chain with a depth cap, from a trigger handler or a controller, after your own DML has finished:
// Written against API version 67.0 (Summer '26)
public with sharing class OrderSyncStarter {
public static Id start(Id orderId) {
AsyncOptions options = new AsyncOptions();
options.MaximumQueueableStackDepth = 4; // first job is depth 1, as in the Queueable Apex stack-depth sample
return System.enqueueJob(new OrderSyncJob(orderId), options);
}
}
The test. Apex test methods cannot make real callouts, so the mock scripts the responses, including a thrown CalloutException for the timeout case.
// Written against API version 67.0 (Summer '26)
@IsTest
private class OrderSyncJobTest {
private class ScriptedMock implements HttpCalloutMock {
private final Integer statusCode;
private final Boolean throwTimeout;
ScriptedMock(Integer statusCode, Boolean throwTimeout) {
this.statusCode = statusCode;
this.throwTimeout = throwTimeout;
}
public HttpResponse respond(HttpRequest req) {
if (throwTimeout) {
throw new System.CalloutException('Read timed out');
}
HttpResponse res = new HttpResponse();
res.setStatusCode(statusCode);
if (statusCode == 429) {
res.setHeader('Retry-After', '90');
}
return res;
}
}
private static Id runJob(Integer statusCode, Boolean throwTimeout) {
Account record = new Account(Name = 'Callout test');
insert record;
Test.setMock(HttpCalloutMock.class, new ScriptedMock(statusCode, throwTimeout));
Test.startTest();
System.enqueueJob(new OrderSyncJob(record.Id));
Test.stopTest();
return record.Id;
}
@IsTest
static void timeoutSchedulesFirstRetryAfterOneMinute() {
runJob(null, true);
Assert.areEqual(1, OrderSyncJob.lastScheduledDelayMinutes);
Assert.areEqual(0, [SELECT COUNT() FROM Callout_Failure__c]);
}
@IsTest
static void rateLimitHonoursRetryAfter() {
runJob(429, false);
Assert.areEqual(2, OrderSyncJob.lastScheduledDelayMinutes); // 90 seconds rounds up to 2 minutes
}
@IsTest
static void serverErrorIsRetriedNotParked() {
runJob(503, false);
Assert.areEqual(1, OrderSyncJob.lastScheduledDelayMinutes);
Assert.areEqual(0, [SELECT COUNT() FROM Callout_Failure__c]);
}
@IsTest
static void clientErrorIsParked() {
Id recordId = runJob(400, false);
Callout_Failure__c parked = [SELECT Record_Id__c, Status_Code__c FROM Callout_Failure__c];
Assert.areEqual(String.valueOf(recordId), parked.Record_Id__c);
Assert.areEqual('400', parked.Status_Code__c);
}
}
Best Practice: If a failure must be handled even when the Queueable dies with an unhandled exception, attach a Transaction Finalizer. The Apex Developer Guide says the finalizer runs in its own transaction and can include REST callouts, and its LoggingFinalizer sample commits a log after the Queueable hits a limit error. A job that failed with an unhandled exception can be re-enqueued from a finalizer up to five times.
How do you debug it when it happens in production?
Work through these checks in order, and write down the answer to each before moving on. The first three usually find it.
- Read the exact exception text and match it to the list above. Timeout, limit, transaction order and endpoint registration are different problems with different owners.
- Compare the timeout to the endpoint’s real latency. Ask the downstream team for their p99 over the failure window. If it is above your
setTimeoutvalue, or above the 10-second default you never changed, you have found it. - Count callouts in the transaction.
Limits.getCallouts()andLimits.getLimitCallouts()in a debug statement before the loop tell you how close you are to 100. - Turn on the Callout debug log category and read the
CALLOUT_REQUESTandCALLOUT_RESPONSElines. Salesforce Help notes these lines survive log truncation, so they are there even in a large log. - Check endpoint coverage. Remote Site Setting for the exact host, or a Named Credential, which the Apex Developer Guide says lets you skip Remote Site Settings for that site.
- Check the concurrent long-running request limit. If failures cluster at peak and the message is
Unable to Process RequestwithConcurrent requests limit exceeded, the problem is long synchronous requests, not the endpoint. Check whether callouts are among them with theApexExecutionandConcurrentLongRunningApexLimitevent types inEventLogFile. The documented count excludes callout wait. One 2020 practitioner report found synchronous callout time counting anyway, confirmed by Salesforce support and by Apex Execution event logs showing millisecond CPU time against callout time over 5 seconds, so measure rather than assume. - Use the Apex Callout event type in
EventLogFilefor a per-callout history across the org. In Enterprise, Unlimited and Performance editions the paid event types need Salesforce Shield or the Event Monitoring add-on, and the EventLogFile Supported Event Types page says which types those are. Developer Edition gets every type, kept for one day. - Correlate by ID. Give the downstream team the
X-Correlation-Idvalues from your failures and ask what their logs show at those moments: slow, rejected, or never received.
The question to settle before you change code is whether the downstream system saw the request. “Never received” is network or configuration. “Received and slow” is timeout tuning or an async path. “Received and rejected” is a payload or authentication problem no retry will fix.
What changes at ten times the volume?
The ceiling moves from the endpoint to the platform, and the design has to stop making callouts on the synchronous path. Three things change:
- Concurrent long-running requests become the limit. The governor limits page excludes HTTP callout time from the 5-second count, but everything around a synchronous callout still runs on the clock: the SOQL before it, the serialising, the DML and trigger work after it. The practitioner report above found the wait itself counting in practice. At ten times the load, as few as ten long requests at once lock out every other Apex request in an org at the minimum limit, the failure Salesforce’s 2015 engineering blog describes. Move callouts off triggers and controllers into Queueables, or use Continuations for user-facing calls.
- Per-endpoint budgets replace one global timeout. Each endpoint gets its own timeout and retry ceiling, held in Custom Metadata, set from its own measured latency. One slow partner no longer sets the budget for the fast ones.
- A circuit breaker stops the retry storm. After a configured number of consecutive transient failures to one endpoint, record an open state with a time to try again, and have the classifier return fail without calling out until then. Platform Cache or a custom setting holds the state. Without it, every retry during an outage spends callouts, async executions and the downstream’s recovery capacity.
The daily asynchronous Apex execution limit also starts to matter: a retry design with four attempts per record multiplies your async usage by up to four during an outage. The circuit breaker is what keeps that bounded.
References
- Apex Developer Guide, Execution Governors and Limits and the Apex Governor Limits cheat sheet: 100 callouts per transaction, 120-second cumulative callout timeout, the concurrent long-running request limit (one per 100 licences, minimum 10, maximum 50) and its exclusion of HTTP callout processing time; the callout limits are the same in the synchronous and asynchronous columns.
- Apex Developer Guide, Callout Limits and Limitations: the 10-second default, the 1 millisecond minimum and 120,000 millisecond maximum, and the 120-second cumulative timeout per transaction.
- Apex Reference Guide, HttpRequest Class (
setTimeout, 1 to 120,000 milliseconds) and HttpResponse Class. - Salesforce Help, Salesforce Apex Callout Timeout Error to Third-Party REST API: a timeout surfaces as
System.CalloutException; the 10-second default and the 1 to 120,000 millisecond range ofsetTimeout. - Salesforce Help, WebService Error: You have uncommitted work pending.
- Metadata API Developer Guide, RemoteSiteSetting: a site must be registered before an Apex callout can call it.
- Apex Developer Guide, Named Credentials as Callout Endpoints.
- Apex Developer Guide, Queueable Apex: one child job per job, delays of 0 to 10 minutes, the delay ignored in tests, chained jobs testable with a stack depth, the default depth of five in Developer and Trial Editions; Apex Reference Guide, AsyncOptions Class and AsyncInfo Class.
- Salesforce Developers Blog, Write Simplified and Secure Apex with Spring ‘23 Updates: the Queueable delay and its rate-limiting use case.
- Apex Developer Guide, Transaction Finalizers, and Salesforce Developers Blog, Introducing Transaction Finalizers: the finalizer runs in its own transaction and can include REST callouts, its LoggingFinalizer sample commits a log after a limit error, and a job that failed with an unhandled exception can be re-enqueued up to five times.
- Apex Reference Guide, Limits Class:
getCallouts,getLimitCallouts; Exception Class and Built-In Exceptions:LimitExceptioncannot be caught,CalloutException; UUID Class. - Lightning Web Components Developer Guide, Continuations: long-running callouts off the synchronous request.
- Salesforce Help, Debug Logs:
CALLOUT_REQUESTandCALLOUT_RESPONSElines are kept when a log is truncated. - Object Reference, Apex Callout Event Type and Concurrent Long-Running Apex Limit Event Type; Salesforce Help, Event Monitoring FAQ: in Enterprise, Unlimited and Performance editions the free version has limited event types and the add-on gives all of them, while Developer Edition has all types for one day; Object Reference, EventLogFile Supported Event Types lists which types are paid.
- Salesforce Engineering Blog, Avoiding Apex Speeding Tickets: synchronous callouts and the concurrent request limit,
Concurrent requests limit exceeded. - Salesforce Developers Blog, Analyze Concurrent Errors in Scale Center: the limit is set by licence type and count, minimum 10, capped at 50.
- Luke Freeland, Salesforce Concurrent Long Running Apex Limit Troubleshooting, October 2020: the practitioner report that synchronous callout time counted in practice, confirmed by Salesforce support and by Apex Execution event logs showing millisecond CPU time against callout time over 5 seconds.
- Trailhead, Apex REST Callouts: status codes are checked by your code; test methods cannot make callouts, use
HttpCalloutMock. - Salesforce Developers Blog, The Salesforce Developer’s Guide to the Summer ‘26 Release: API version 67.0.
- IETF, RFC 6585 (status 429), RFC 9110 (
Retry-After), and the HTTP API working group’s Idempotency-Key header Internet-Draft, expired in its current revision. - Stripe API Reference, Idempotent requests: the
Idempotency-Keyheader in production use.
Need an integration that fails less and tells you why?
I take short specialist contracts to diagnose intermittent Salesforce integration failures and rebuild the retry, timeout and observability layer around them.