Error body format and reading order
All failure responses share the same structure:
{"error":{"code":"...","message":"..."}}
Reading order is fixed in three steps:
- HTTP status code determines the category;
error.codedetermines the specific cause; use it for branching in code;error.messageis for humans; log it, but don't use it for string matching.
First, define the boundary: things solvable by retrying (429, 503, network timeouts) and things that must be changed to solve (the rest). Mixing these two is the most common root cause of production incidents, such as retrying 402 endlessly and sending dozens of invalid requests per second.
Log four fields consistently: timestamp, status code, error.code, and estimated prompt_tokens for the request. This instantly shows if a specific error type spiked or if overall failures increased. Set an alert: notify immediately on any 402 (service is unavailable to the user); for 429, monitor the ratio—sporadic is fine, persistent indicates concurrency design issues.
Status code quick reference
| Status code | error.code | Meaning | Retry? |
|---|---|---|---|
| 400 | — | Invalid request, e.g., prompt + max_tokens exceeds 100k | No |
| 401 | — | Invalid or missing key | No |
| 402 | no_credit | Balance exhausted or trial expired | No |
| 403 | content_blocked | Content blocked | No |
| 404 | — | Endpoint not found | No |
| 429 | — | Rate limit exceeded | Yes, back off |
| 503 | upstream_busy | Service temporarily busy | Yes, after a few seconds |
A dash in the table means there is no fixed code string you need to handle specifically; branch on the status code instead.
4xx: Fix the request, don't retry
400 Invalid request
- Cause 1: JSON format error, often due to unescaped quotes when hand-crafting strings.
- Cause 2: prompt + max_tokens exceeds 100,000. Long conversations are easiest to trigger.
- Cause 3: Request body exceeds 8 MB.
- Fix: Serialize with a JSON library; estimate tokens before sending; trim history or reduce max_tokens if over limit; see Long Context in Practice.
401 Key issues
- Missing Header or missing
Bearerprefix. - Regenerated the key, invalidating the old one immediately, while some machines still use the old value.
- Environment variable not passed into the container or scheduled task.
402 no_credit
Balance is empty or the 7-day trial has expired. Top up prepaid credit on the account page to restore access. Monitor your balance so you don't wait for users to report errors.
403 content_blocked
Content blocked. Legitimate adult content, fiction, and controversial topics are not rejected, but sexual content involving minors is always blocked, including in novels and roleplay. If you get a 403, check the input and history for such content instead of rephrasing and retrying.
404 endpoint not found
Only two endpoints: POST /v1/chat/completions and GET /v1/models. Missing /v1, extra slashes, typos, or requesting unsupported embeddings or image endpoints will 404.
Another easily misdiagnosed 400 case: token count grows linearly with turns. Testing works fine during the day, but suddenly 400s appear after dozens of turns. This is not instability; the context window is full. Cap history, drop oldest turns when exceeding the threshold, or compress old content into summaries. Estimate before sending requests.
A tip for troubleshooting 401: log the first and last four characters of the API key, not the full value, and compare it with what the account page shows. This confirms whether the process actually read the key you think it did. In container and cron environments, missing environment variables or reading stale values are the most common causes.
429 and 503: the two classes that require retrying
429 rate limit
300 requests per key per minute. Batch tasks, multiple instances sharing one key, and retry storms will all trigger this. Fixes: client-side rate limiting (see code below), then exponential backoff retry for 429.
503 upstream_busy
Service is busy. Retry after a few seconds. Don't blast requests immediately or retry ten times in one second; that will only make things worse.
Network timeout
Long text generation takes time; set timeout to 120s. When retrying after timeout, note that the previous request may have already executed on the server, so duplicate calls may consume quota twice. For generation requests, limit retries and use streaming to shorten wait times.
import os
import random
import time
import requests
URL = "https://api.wushenchaapi.com/v1/chat/completions"
HEADERS = {
"Authorization": "Bearer " + os.environ["API_KEY"],
"Content-Type": "application/json",
}
RETRYABLE = {429, 503} # 只重试这两类
FATAL = {400, 401, 402, 403, 404}
class ApiError(Exception):
def __init__(self, status, code, message):
super().__init__(f"{status} {code}: {message}")
self.status, self.code = status, code
def call(payload, max_retries=5, base=1.0, cap=30.0):
for attempt in range(max_retries + 1):
try:
r = requests.post(URL, headers=HEADERS, json=payload, timeout=120)
except (requests.ConnectionError, requests.Timeout):
r = None # 网络层失败,按可重试处理
if r is not None and r.ok:
return r.json()
if r is not None:
try:
err = r.json().get("error", {})
except ValueError:
err = {}
if r.status_code in FATAL or r.status_code not in RETRYABLE:
raise ApiError(r.status_code, err.get("code"), err.get("message"))
if attempt == max_retries:
raise ApiError(r.status_code if r is not None else 0, "retry_exhausted", "重试次数用尽")
delay = min(cap, base * (2 ** attempt)) * random.uniform(0.5, 1.0)
time.sleep(delay)
if __name__ == "__main__":
out = call({"model": "uncensored", "max_tokens": 100,
"messages": [{"role": "user", "content": "回复一个字:好"}]})
print(out["choices"][0]["message"]["content"])Key: backoff formula is min(cap, base × 2^n) × random factor. Random jitter avoids multiple clients retrying simultaneously and colliding. Never retry FATAL status codes.
Client-side rate limiting and diagnostic requests
Instead of waiting for 429 to back off, limit requests beforehand. The tool below guarantees request count stays below the set value per minute and is thread-safe:
import threading
import time
class RateGate:
"""简单的客户端限速:保证一分钟内请求数不超过 limit。"""
def __init__(self, limit=240): # 留出余量,低于 300/分钟
self.limit, self.stamps, self.lock = limit, [], threading.Lock()
def wait(self):
while True:
with self.lock:
now = time.time()
self.stamps = [t for t in self.stamps if now - t < 60]
if len(self.stamps) < self.limit:
self.stamps.append(now)
return
sleep_for = 60 - (now - self.stamps[0])
time.sleep(max(sleep_for, 0.05))For troubleshooting, the cleanest method is to bypass business code and send a minimal request with curl. -i lets you see both the status line and response headers:
curl -i https://api.wushenchaapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"uncensored","max_tokens":20,"messages":[{"role":"user","content":"ping"}]}'If curl works but your business code fails, the issue is in your code or environment (proxy, environment variables, encoding). If curl also fails, check the API key, balance, and network.
Why set the rate limit to 240 instead of 300: in multi-instance deployments, each instance limits itself, but the sum may still exceed the total. Retries also consume extra quota. Leaving a 20% buffer is a safer approach. If you have multiple instances, divide the total quota equally among them, or use a shared counter for unified scheduling.
Troubleshooting checklist
- Read the
error.codein the body; do not rely solely on the status code. - 401: API key prefix, environment variables, whether it was regenerated.
- 400: Is JSON valid? Does prompt + max_tokens exceed 100,000? Does request body exceed 8 MB?
- 402: Balance and free trial validity.
- 403: Does the input or history contain prohibited content?
- 404: Is the path /v1/chat/completions or /v1/models?
- 429: Are multiple instances sharing one API key? Is client-side rate limiting implemented?
- 503: Is backoff retry implemented? Is the interval at least a few seconds?
- Timeout: Is the timeout sufficient? Can you switch to streaming?
- If all else is ruled out, reproduce the issue with a minimal curl request.
If just getting started, read the integration guide. More parameters are in the documentation.
How to use the checklist: eliminate issues top-to-bottom when problems occur; do not skip. Most "weird" bugs fall into the first four items.
Pro tip: write troubleshooting conclusions back to team docs. Next time the same error appears, you can match it directly to the solution.