Rate Limits & Errors
Limits
Sapling's API is rate-limited based on the subscription of the API key holder.
- For free trials, the limit is 50,000 characters every 24 hours and 250,000 characters/month.
Free-trial accounts also have a daily character allowance shared across all
endpoints. When it runs out, requests return
402until the next UTC midnight. - For subscribed users, this limit is currently 10 million characters every 24 hours and 25 million characters/month.
- Each of these limits can be raised – contact us.
How the limits are applied:
- Per key and per endpoint. Each endpoint keeps its own allowance for each API key, so heavy use of one endpoint doesn't slow down another.
- Continuous refill. The allowance refills evenly over its window instead of resetting all at once. For example, a limit of 50,000 characters every 24 hours gets back roughly 35 characters per minute.
- Batches. A batch request (
texts) counts the combined length of its items. - Model-backed endpoints. Classify, extract, named entities, simplify, style guide, summarize and translate count a minimum number of characters per request, and a batch counts that minimum once per item. Many very short requests therefore use more allowance than their length alone.
- Monthly quota. The monthly figure is a per-key quota that resets at the start of each UTC month.
To see where a key stands, including why it is blocked and for how long, call Account Status.
Rate-limit headers
Rate-limited endpoints report the state of your allowance on each response they
serve, including 429s:
| Header | Meaning |
|---|---|
X-RateLimit-Limit | Size of this endpoint's allowance for your key, in characters. |
X-RateLimit-Remaining | Characters currently available. |
X-RateLimit-Reset | Seconds until the allowance is completely full again. This is a number of seconds, not a timestamp. |
Use these headers to pace a batch job before you hit a 429. They are omitted
for keys that are configured without a rate limit. They are also omitted on
self-hosted deployments that turn API rate limiting off.
Error format
Errors are JSON objects with a human-readable msg field:
{"msg": "Invalid API key."}
This applies to request-validation errors, authentication errors, unknown paths, wrong HTTP methods and unexpected server errors. A few responses differ:
- A request body over 4MB (
413) usesmessageandresultkeys instead ofmsg. /api/v1/editscan answer a temporary internal outage with an empty JSON array ([]) and status400.- Request-validation errors prefix
msgwith400 Bad Request:, for example"400 Bad Request: Missing text argument".
Treat the status code as authoritative and use msg for logs and messages
to users. The OpenAPI specification lists the exact statuses
each endpoint can return.
Status codes
| Status | Meaning | Retry? |
|---|---|---|
400 | The request is invalid: malformed JSON, or a missing or invalid argument. | No. Fix the request. Exception: the safety, language detection, AI detector, rephrase, tone and sentiment endpoints also answer a temporary model failure with a 400 whose msg starts with Unexpected error. That case can be retried after a short delay. |
401 | Missing or invalid API key. | No. |
402 | Payment required. See below. | Depends on the reason. |
403 | The request isn't allowed for this key. For example, a team_id names a team the key's user doesn't belong to, or an MCP OAuth token was used on an endpoint outside the MCP tool set. | No. |
404 | Unknown path, or a named resource (such as a saved style-guide ruleset) doesn't exist. | No. |
405 | Wrong HTTP method. The Allow header lists the methods the path supports. | No. |
413 | Request body over 4MB. | No. Split the text. |
429 | Rate limited. See below. | Usually, after waiting. |
500 | Unexpected server error. | Once or twice with backoff. Report persistent errors to support@sapling.ai. |
502 | A model-backed endpoint's upstream model call failed. Nothing is cached or billed. | Yes, after a short delay. |
503 | Temporarily overloaded. | Yes. Wait for the Retry-After interval. |
402 Payment Required
A 402 has one of two meanings:
- The block is permanent until you act. The key or subscription has expired,
prepaid credits are used up, or the account reached its monthly billing limit.
Retrying won't help, so fix the account in your
API settings. The response carries
Retry-After: 3600only as a hint not to retry in a tight loop. See payment-required responses. - The block is temporary. A free-trial account used up its daily allowance.
The block clears at the next UTC midnight. Both
Retry-Afterand a numericretry_afterbody field give the seconds until then. The presence ofretry_afteris what tells the two cases apart.
429 Rate Limited
A 429 also comes in a few forms. Check the body's retry_after field:
retry_afteris present. You used up your allowance for the moment. Wait that many seconds, which is also sent as theRetry-Afterheader, then retry. This covers two cases:- The short-term allowance is used up.
msgstarts withRate limited. Capacity used. - The key's monthly quota is used up.
msgstarts withRate Limited. Monthly quota used., andretry_aftercounts down to the start of the next UTC month.
- The short-term allowance is used up.
retry_afteris absent. This request can never succeed as sent, because it is larger than the key's entire allowance. It also happens when a request without an API key goes over the free character cap on rephrase or summarize. Don't retry unchanged. Split the text or send an API key.- File uploads. The file endpoints allow 60 uploads per minute and send no
Retry-Afteror rate-limit headers. Back off for a minute and retry.
Retrying requests
A simple, safe policy:
- Retry these responses:
- a
429or402whose body includesretry_after(a short-term limit, the monthly quota, or the free-trial daily allowance) 500,502and503responses- the
400 Unexpected errorcase described above
- a
- Wait at least
Retry-Afterseconds when it's present. Otherwise use exponential backoff starting at about one second. - Give up after a few attempts, or when the wait is longer than your job can
tolerate. A trial
402can mean waiting until UTC midnight, and a monthly-quota429until the next month. Never retry a429or402withoutretry_afterin a loop. - File uploads are the exception. Their
429never hasretry_after, so back off for about a minute and retry.
The helper below is for the JSON endpoints. It gives up rather than sleep longer
than max_wait seconds.
import time
import requests
RETRYABLE = {500, 502, 503}
def post_with_retries(url, payload, attempts=4, max_wait=300):
delay = 1
for attempt in range(attempts):
response = requests.post(url, json=payload)
try:
body = response.json() if response.headers.get('Content-Type', '').lower().startswith('application/json') else {}
except ValueError:
body = {}
if response.status_code < 400:
return body
transient_400 = (response.status_code == 400 and isinstance(body, dict)
and str(body.get('msg', '')).startswith('Unexpected error'))
can_wait = (response.status_code in (402, 429) and isinstance(body, dict)
and 'retry_after' in body)
if attempt == attempts - 1 or not (response.status_code in RETRYABLE or transient_400 or can_wait):
response.raise_for_status()
try:
sleep_time = float(response.headers.get('Retry-After', delay))
except ValueError:
sleep_time = delay
if sleep_time > max_wait:
response.raise_for_status()
time.sleep(sleep_time)
delay *= 2
If you keep running into rate limits, contact support@sapling.ai and we'll see
if we can help.