Calculate limits first
Three numbers determine everything:
- Total context is 100,000 tokens, including both prompt and completion.
max_tokensdefaults to 2048, with a maximum of 16,000 per request.- The request body must not exceed 8 MB; this limit is rarely hit for plain text, whereas tokens hit the cap first.
The budget formula is simple: prompt available = 100000 - max_tokens. If you plan for the model to output 4,000 tokens at once, the prompt is left with only 60,000. When exceeded, the interface returns 400; it won't truncate automatically. So "stuffing the whole book in first" is not a strategy, it's just luck.
Another common misconception: setting max_tokens higher is safer. It actually reserves output space; setting it to 16,000 means your prompt is limited to 48,000. Set it as needed, don't just max it out.
Example: assume you need to summarize a 300,000-character interview transcript. At ~1.3 tokens per character, that's ~390,000 tokens, more than six times the limit, so you can't send it all in. If each summary output is 800 tokens, each input chunk can be at most 63,000. Leaving a 10% buffer, that's about 50,000. But this isn't optimal; larger chunks reduce the model's focus on middle content. The better approach is smaller chunks, even if it means more calls.
Token estimation: approximate algorithm
Without a local tokenizer, use these values for estimation. These are approximations, not exact values, with variance across text types reaching 10-20%:
| Text type | Approximate conversion |
|---|---|
| Chinese prose | Approx. 1 to 1.5 tokens per character |
| English | Approx. 1 token per 4 characters, or ~1.3 tokens per word |
| Code, JSON, symbol-heavy text | Consumes more tokens than normal text; estimate using a conservative high value |
A sufficient estimation function with budgeting for buffer:
CONTEXT = 100000
def est_tokens(text: str) -> int:
"""粗估:中文按每字 1.3 token,其余字符按每 4 个字符 1 token。仅为近似。"""
zh = sum(1 for c in text if "\u4e00" <= c <= "\u9fff")
other = len(text) - zh
return int(zh * 1.3 + other / 4) + 1
def max_prompt_budget(max_tokens: int, margin: float = 0.1) -> int:
"""给定计划的 max_tokens,返回 prompt 最多能用多少 token(留 margin 余量)。"""
return int((CONTEXT - max_tokens) * (1 - margin))
# 例:计划让模型写 3000 token,prompt 预算
print(max_prompt_budget(3000)) # 54900Calibration: Send a typical input request, read usage.prompt_tokens from the response, compare with your estimate, calculate the real multiplier for your data, and update the coefficient. Estimates decide "whether to chunk"; check usage for billing.
Long document summarization: chunk and merge
When a document exceeds the budget, the standard approach is map-reduce: chunk and summarize each part, then merge summaries into a total summary. Note the following details:
- Split at paragraph boundaries, not by fixed character counts, to avoid cutting sentences in half.
- Leave sufficient budget per chunk; recommend keeping each chunk under one-third of the limit to leave room for instructions and output.
- Chunk summaries should preserve names, numbers, and key events, otherwise information is lost during merging.
- If the summary of summaries is still too long, recurse one more level.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.wushenchaapi.com/v1", api_key=os.environ["API_KEY"])
def split_paragraphs(text, budget):
"""按段落切块,每块估算不超过 budget token。"""
chunks, cur, used = [], [], 0
for p in text.split("\n"):
t = int(len(p) * 1.3) + 1 # 中文粗估
if used + t > budget and cur:
chunks.append("\n".join(cur))
cur, used = [], 0
cur.append(p)
used += t
if cur:
chunks.append("\n".join(cur))
return chunks
def ask(prompt, max_tokens=800):
r = client.chat.completions.create(
model="uncensored",
temperature=0.3,
max_tokens=max_tokens,
messages=[{"role": "user", "content": prompt}],
)
return r.choices[0].message.content
def summarize_long(text, budget=12000):
parts = split_paragraphs(text, budget)
partial = [ask("用不超过 200 字概括下面这段,保留人名和关键事件:\n\n" + p) for p in parts]
merged = "\n".join(f"第{i+1}部分:{s}" for i, s in enumerate(partial))
return ask("下面是分段摘要,请合并成一份 400 字以内的整体摘要:\n\n" + merged, max_tokens=900)Set temperature to 0.3 for stable summarization. Calls are independent and can be parallel requests, but watch the rate limit of 300 requests per key per minute; a document of dozens of chunks won't hit this.
There is no standard answer for chunk size. Experience suggests 10,000 to 15,000 tokens per chunk balances detail retention and call frequency. Test a sample with chunk sizes of 5,000, 12,000, and 25,000, and compare the number of key facts missed in the summaries to choose the best value for your document type. Contracts and technical docs (high information density) need smaller chunks; dialogue logs (high redundancy) can use larger chunks.
Novel continuation: sliding window and rolling outline
By the tenth chapter, the full text won't fit and doesn't need to. The approach uses two layers of memory:
- Short-term:Last 2-3 chapters to maintain style, dialogue rhythm, and scene continuity.
- Long-term:Earlier chapters compressed into an outline, preserving character relations, foreshadowing, and unresolved clues.
After writing each chapter, have the model summarize it into three or five lines and append it to the outline. If the outline exceeds 3000 tokens, compress it as a whole.
def continue_story(chapters, outline, new_hint, keep_last=3):
"""chapters: 已写章节列表。只带最近 keep_last 章原文,更早的用 outline(滚动大纲)代替。"""
recent = "\n\n".join(chapters[-keep_last:])
prompt = (
f"【全书大纲(早期章节摘要)】\n{outline}\n\n"
f"【最近章节原文】\n{recent}\n\n"
f"【下一章要求】\n{new_hint}\n\n请写下一章,约 2000 字。"
)
return ask(prompt, max_tokens=4000)Foreshadowing is the easiest thing to lose. Add a "unresolved foreshadowing" section to the outline and explicitly require the model to address one item per chapter. List character and place names in a fixed list in the system prompt to prevent the model from renaming them when they slide out of the window.
Another often-overlooked issue in continuation is style drift. The model tends to mimic the style of the most recent chapters, amplifying inconsistencies. Fix this by adding a "style sample" to the system prompt: paste a paragraph of your best original text as a fixed anchor that doesn't slide with the window.
max_tokens and truncation
There are two scenarios for truncated responses, which you must distinguish:
finish_reasonislength: you hit the max_tokens limit. Fix by increasing max_tokens or letting the model write in chunks: "stop here, I say continue then write."- Request returns 400 directly: prompt plus max_tokens exceeds 100,000. Fix by shortening input or reducing max_tokens.
A practical budget reference:
| Task | Suggested max_tokens | prompt available (approx.) |
|---|---|---|
| summarization, extraction | 800 | 63,200 |
| single-chapter continuation (approx. 2,000 words) | 4000 | 60,000 |
| 10,000-word long output | 16,000 (limit) | 48,000 |
Long outputs are also affected by response time; use streaming to display as it generates.
When continuation is truncated, do not simply feed the partial content to the model and say 'continue.' A more robust approach is to include the generated portion as an assistant message in messages, then add a user message saying 'continue from the last sentence, do not repeat.' This ensures the smoothest transition. Note that this step increases the prompt length, so recalculate your budget.
Long context checklist
- Estimate the prompt size before each request; if it exceeds the budget, use the chunking branch.
- Record usage for every response to continuously calibrate estimation coefficients.
- Place the system prompt and outline at the beginning; reiterate key instructions at the end.
- Check finish_reason; if it is length, trigger an alert or continuation.
- Set a retention limit for conversation history; discard the oldest messages when the limit is reached instead of waiting for the API to return a 400 error.
Related: Prompt structure in Prompt Writing, error handling in Error Code Troubleshooting, pricing in Price page.
Final reminder: long context does not equal long-term memory. The model reads your content from scratch for each request and does not remember anything from previous calls. You must provide the history yourself to maintain continuity.