EN ▾
Get API key

100k long context in practice: continuation, summarization, chunking, budgeting

A 100,000 token context window sounds generous, but in practice: a full novel won't fit in, and a dozens-page document plus required output will hit the limit. This treats the long context as a budget: how to estimate tokens, how to reserve space for output, how to chunk and summarize long texts, and how to use a sliding window with a rolling outline for novel continuation. Code in Python, logic applies to any language.

Updated on

Key points

  • 100,000 is the sum of prompt and completion, not a separate input limit.
  • Chinese token count is only an estimate, roughly 1 to 1.5 per character; rely on usage in the response for accuracy.
  • Summarize long texts by chunking and merging; continue novels using recent chapters and a rolling outline.
  • max_tokens defaults to 2048 and maxes out at 16,000; calculate your budget before setting it for long content.

Calculate limits first

Three numbers determine everything:

  • Total context is 100,000 tokens, including both prompt and completion.
  • max_tokens defaults to 2048, with a maximum of 16,000 per request.
  • The request body must not exceed 8 MB; this limit is rarely hit for plain text, whereas tokens hit the cap first.

The budget formula is simple: prompt available = 100000 - max_tokens. If you plan for the model to output 4,000 tokens at once, the prompt is left with only 60,000. When exceeded, the interface returns 400; it won't truncate automatically. So "stuffing the whole book in first" is not a strategy, it's just luck.

Another common misconception: setting max_tokens higher is safer. It actually reserves output space; setting it to 16,000 means your prompt is limited to 48,000. Set it as needed, don't just max it out.

Example: assume you need to summarize a 300,000-character interview transcript. At ~1.3 tokens per character, that's ~390,000 tokens, more than six times the limit, so you can't send it all in. If each summary output is 800 tokens, each input chunk can be at most 63,000. Leaving a 10% buffer, that's about 50,000. But this isn't optimal; larger chunks reduce the model's focus on middle content. The better approach is smaller chunks, even if it means more calls.

Token estimation: approximate algorithm

Without a local tokenizer, use these values for estimation. These are approximations, not exact values, with variance across text types reaching 10-20%:

Text typeApproximate conversion
Chinese proseApprox. 1 to 1.5 tokens per character
EnglishApprox. 1 token per 4 characters, or ~1.3 tokens per word
Code, JSON, symbol-heavy textConsumes more tokens than normal text; estimate using a conservative high value

A sufficient estimation function with budgeting for buffer:

CONTEXT = 100000

def est_tokens(text: str) -> int:
    """粗估:中文按每字 1.3 token,其余字符按每 4 个字符 1 token。仅为近似。"""
    zh = sum(1 for c in text if "\u4e00" <= c <= "\u9fff")
    other = len(text) - zh
    return int(zh * 1.3 + other / 4) + 1

def max_prompt_budget(max_tokens: int, margin: float = 0.1) -> int:
    """给定计划的 max_tokens,返回 prompt 最多能用多少 token(留 margin 余量)。"""
    return int((CONTEXT - max_tokens) * (1 - margin))

# 例:计划让模型写 3000 token,prompt 预算
print(max_prompt_budget(3000))   # 54900

Calibration: Send a typical input request, read usage.prompt_tokens from the response, compare with your estimate, calculate the real multiplier for your data, and update the coefficient. Estimates decide "whether to chunk"; check usage for billing.

Long document summarization: chunk and merge

When a document exceeds the budget, the standard approach is map-reduce: chunk and summarize each part, then merge summaries into a total summary. Note the following details:

  1. Split at paragraph boundaries, not by fixed character counts, to avoid cutting sentences in half.
  2. Leave sufficient budget per chunk; recommend keeping each chunk under one-third of the limit to leave room for instructions and output.
  3. Chunk summaries should preserve names, numbers, and key events, otherwise information is lost during merging.
  4. If the summary of summaries is still too long, recurse one more level.
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.wushenchaapi.com/v1", api_key=os.environ["API_KEY"])

def split_paragraphs(text, budget):
    """按段落切块,每块估算不超过 budget token。"""
    chunks, cur, used = [], [], 0
    for p in text.split("\n"):
        t = int(len(p) * 1.3) + 1          # 中文粗估
        if used + t > budget and cur:
            chunks.append("\n".join(cur))
            cur, used = [], 0
        cur.append(p)
        used += t
    if cur:
        chunks.append("\n".join(cur))
    return chunks

def ask(prompt, max_tokens=800):
    r = client.chat.completions.create(
        model="uncensored",
        temperature=0.3,
        max_tokens=max_tokens,
        messages=[{"role": "user", "content": prompt}],
    )
    return r.choices[0].message.content

def summarize_long(text, budget=12000):
    parts = split_paragraphs(text, budget)
    partial = [ask("用不超过 200 字概括下面这段,保留人名和关键事件:\n\n" + p) for p in parts]
    merged = "\n".join(f"第{i+1}部分:{s}" for i, s in enumerate(partial))
    return ask("下面是分段摘要,请合并成一份 400 字以内的整体摘要:\n\n" + merged, max_tokens=900)

Set temperature to 0.3 for stable summarization. Calls are independent and can be parallel requests, but watch the rate limit of 300 requests per key per minute; a document of dozens of chunks won't hit this.

There is no standard answer for chunk size. Experience suggests 10,000 to 15,000 tokens per chunk balances detail retention and call frequency. Test a sample with chunk sizes of 5,000, 12,000, and 25,000, and compare the number of key facts missed in the summaries to choose the best value for your document type. Contracts and technical docs (high information density) need smaller chunks; dialogue logs (high redundancy) can use larger chunks.

Novel continuation: sliding window and rolling outline

By the tenth chapter, the full text won't fit and doesn't need to. The approach uses two layers of memory:

  • Short-term:Last 2-3 chapters to maintain style, dialogue rhythm, and scene continuity.
  • Long-term:Earlier chapters compressed into an outline, preserving character relations, foreshadowing, and unresolved clues.

After writing each chapter, have the model summarize it into three or five lines and append it to the outline. If the outline exceeds 3000 tokens, compress it as a whole.

def continue_story(chapters, outline, new_hint, keep_last=3):
    """chapters: 已写章节列表。只带最近 keep_last 章原文,更早的用 outline(滚动大纲)代替。"""
    recent = "\n\n".join(chapters[-keep_last:])
    prompt = (
        f"【全书大纲(早期章节摘要)】\n{outline}\n\n"
        f"【最近章节原文】\n{recent}\n\n"
        f"【下一章要求】\n{new_hint}\n\n请写下一章,约 2000 字。"
    )
    return ask(prompt, max_tokens=4000)

Foreshadowing is the easiest thing to lose. Add a "unresolved foreshadowing" section to the outline and explicitly require the model to address one item per chapter. List character and place names in a fixed list in the system prompt to prevent the model from renaming them when they slide out of the window.

Another often-overlooked issue in continuation is style drift. The model tends to mimic the style of the most recent chapters, amplifying inconsistencies. Fix this by adding a "style sample" to the system prompt: paste a paragraph of your best original text as a fixed anchor that doesn't slide with the window.

max_tokens and truncation

There are two scenarios for truncated responses, which you must distinguish:

  1. finish_reason is length: you hit the max_tokens limit. Fix by increasing max_tokens or letting the model write in chunks: "stop here, I say continue then write."
  2. Request returns 400 directly: prompt plus max_tokens exceeds 100,000. Fix by shortening input or reducing max_tokens.

A practical budget reference:

TaskSuggested max_tokensprompt available (approx.)
summarization, extraction80063,200
single-chapter continuation (approx. 2,000 words)400060,000
10,000-word long output16,000 (limit)48,000

Long outputs are also affected by response time; use streaming to display as it generates.

When continuation is truncated, do not simply feed the partial content to the model and say 'continue.' A more robust approach is to include the generated portion as an assistant message in messages, then add a user message saying 'continue from the last sentence, do not repeat.' This ensures the smoothest transition. Note that this step increases the prompt length, so recalculate your budget.

Long context checklist

  • Estimate the prompt size before each request; if it exceeds the budget, use the chunking branch.
  • Record usage for every response to continuously calibrate estimation coefficients.
  • Place the system prompt and outline at the beginning; reiterate key instructions at the end.
  • Check finish_reason; if it is length, trigger an alert or continuation.
  • Set a retention limit for conversation history; discard the oldest messages when the limit is reached instead of waiting for the API to return a 400 error.

Related: Prompt structure in Prompt Writing, error handling in Error Code Troubleshooting, pricing in Price page.

Final reminder: long context does not equal long-term memory. The model reads your content from scratch for each request and does not remember anything from previous calls. You must provide the history yourself to maintain continuity.

Frequently asked questions

Does the 100k context include output?

Yes. The limit is the sum of the prompt and completion, so the larger you set max_tokens, the less input you can fit.

How many tokens is one Chinese character?

You can only estimate roughly: about 1 to 1.5 tokens per character, varying by text. The most accurate method is to read the usage field in the response and calibrate using actual measurements.

Does input get truncated automatically when the limit is exceeded?

No, the request returns 400. You need to chunk or discard old content before calling.

What is the maximum length for a single generation?

max_tokens defaults to 2048, with a maximum of 16,000 per request. For longer content, generate in multiple steps and use an outline or summary to maintain coherence.

Just fill out the form to get your key

Create an account, copy the key, and modify the Base URL. Configuration is that simple.