EN ▾
Get API key

Prompt Writing for Uncensored Model API: system, role, format, sampling

The same model produces vastly different output quality depending on whether the prompt is loose or tight. No magic here, just engineering practices you can verify: how to structure the system prompt, how to define roles without drift, how to get parseable JSON via instructions, how to set temperature and top_p, and which approaches fail. Every point can be tested in your code.

Updated on

Key points

  • Use a two-part system prompt: rules plus setting. Number the rules; state only facts in the setting.
  • Specify the JSON format in the instructions, lower the temperature, and add fallback parsing in your code.
  • Adjust only one parameter at a time. For creative writing, use 0.8 to 1.0; for extraction, start with 0 to 0.3.
  • Common failure cases: vague negations, contradictory rules, and placing format requirements in the middle of the conversation.

Start with four principles

  • One request per task.If you ask the model to write a plot, summarize it, and output JSON all at once, performance drops. Split into multiple requests; the cost is low, at $0.25 per million tokens of input.
  • Rules must be verifiable. "Write more vividly" cannot be verified, but "No more than 120 characters per paragraph, with at least one environmental description" can.
  • Prefer positive phrasing. Saying "Only write actions and dialogue" is much more stable than saying "Do not write psychological activities."
  • Start with small samples, then scale up. For any prompt change, compare with five to ten examples first, then go live.

This model does not refuse legal adult-oriented fictional content, fictional genres, or controversial topics, so you do not need to beat around the bush in the prompt or repeatedly declare "this is just a novel." Writing the task clearly directly is actually more stable. The only hard boundary is sexual content involving minors, which returns 403 regardless of whether it is fictional or not, and this cannot be changed by wording.

Two-part structure for the system prompt

Recommend splitting the system into "rules" and "settings". Put rules first, numbered; put settings last, stating only facts. Here is how a novel narrator writes:

你是「夜航」,一名为成年读者写黑色悬疑小说的叙述者。

# 规则
1. 第三人称过去时,每段不超过 120 字。
2. 不替用户的角色做决定,只写环境和其他角色的反应。
3. 每次输出 300 到 500 字,结尾停在一个未决的动作上。
4. 不总结、不点评、不加免责声明,直接写正文。

# 设定
时间:1998 年深秋。地点:港口城市旧码头区。
主角:沈野,退役水警,嗜烟,右耳有旧伤。

Key points:

  1. No more than six rules. The more rules there are, the more easily the later ones are ignored.
  2. Numeric constraints (paragraph length, output length) are more effective than adjectives.
  3. Ending constraints (stopping on an unresolved action) allow multi-turn continuations to flow naturally.
  4. Do not pack plot progression into the settings. Plot is advanced by user messages; otherwise, the model fights with user input later on.

Match output length to max_tokens. If the rule says 500 characters but max_tokens is 200, the output is truncated mid-sentence.

Another practical detail: numbered rules in the system should ideally express only one thing each. "No more than 120 characters per paragraph, and do not summarize" looks like one rule, but it is two. After splitting, the model can more easily satisfy both. After writing, read each rule and ask: Can this be checked by eye for compliance? If not, rewrite it in a verifiable form.

How to define roles without drift

Character drift is the most common problem in multi-turn conversations: the tone is normal for the first five turns, but by the twentieth turn, it goes out of character. There are three countermeasures.

  • Settings write only observable traits. "Shen Ye: smokes, old injury on the right ear, speaks in short sentences" is better than "Shen Ye is a complex and charming person."
  • Write speaking style as examples. Providing two or three sample lines helps the model mimic the style far better than understanding adjectives.
  • Periodically reiterate. For long conversations, append a short reminder at the end of user messages every few turns, such as "Maintain Shen Ye's short-sentence style," costing only a few dozen tokens.

Also, do not mix character settings with output rules. Rules are "how to write," settings are "who to write." Separating them lets you swap characters without changing rules, making A/B testing easier. For long conversations, watch the total context window size; see Practical Guide to 100k Long Context for details.

Add one more rule for multi-character scenes: set each character on a separate line, give each a sentence of dialogue style as an example, to avoid everyone speaking like the same person. For the user's character, only write "controlled by the user" in the settings, so the model does not write on their behalf.

Controlling JSON output with instructions

Do not assume the model "will" return valid JSON. A reliable approach uses three layers: hardcode the format in instructions, lower randomness via parameters, and add fallback parsing in code.

import json
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.wushenchaapi.com/v1", api_key=os.environ["API_KEY"])

SYSTEM = (
    "你是信息抽取器。只输出一个 JSON 对象,不要 Markdown 代码块,不要任何解释。"
    '格式:{"name": 字符串, "mood": "calm|tense|angry", "items": [字符串]}。'
    "缺失的字段用 null,items 没有则给空数组。"
)

def extract(text, retries=2):
    for _ in range(retries + 1):
        resp = client.chat.completions.create(
            model="uncensored",
            temperature=0.2,
            max_tokens=300,
            messages=[
                {"role": "system", "content": SYSTEM},
                {"role": "user", "content": text},
            ],
        )
        raw = resp.choices[0].message.content.strip()
        raw = raw.removeprefix("```json").removesuffix("```").strip()
        try:
            return json.loads(raw)
        except json.JSONDecodeError:
            continue
    return None

print(extract("老周把钥匙拍在桌上,冷着脸说:账本和那把铜钥匙,今晚都得还我。"))

This code demonstrates several habits:

  • Write format instructions in the system, and give field value ranges (calm|tense|angry).
  • Explicitly prohibit code blocks and explanatory text, while still removing possible fences in the code for a double safety net.
  • Write conventions for missing fields into the prompt (null, empty array) to avoid the model making them up.
  • Retry limit on parse failure; return None at the end, letting the caller decide what to do.

If there are many fields, paste a complete example object in the prompt first, then have it output according to it; this is usually more accurate than describing a bunch of rules.

Recommendations for temperature and top_p

These are starting points, not final conclusions. Base your decision on comparing your own examples. Principle: adjust only one parameter at a time; keep the other at default.

Scenariotemperaturetop_pNotes
JSON extraction, classification0 to 0.3DefaultFor stability; repeated calls should yield consistent results
Rewriting, polishing0.5 to 0.7DefaultPreserves original meaning, allows wording variations
Novel continuation, character dialogue0.8 to 1.00.9 to 0.95High diversity, watch for occasional tangents
Brainstorming, namingAround 1.0DefaultSample multiple times and pick the best

Two signals help you gauge direction: repetitive output and identical sentence structures indicate low temperature; irrelevant content and inconsistent character names indicate high temperature or top_p. The stop parameter is also useful, e.g., stopping the model at a specific marker to facilitate segmented generation.

Common prompt mistakes

PromptIssueFix
"Try not to make it too long"No numbers, impossible to execute"No more than 400 words"
"Don't write A, and don't not write A"Conflicting rulesKeep only one clear rule
Format requirements buried in dialogueDrowned out after long conversationsMove to system, or restate at the end of each turn
Tell the model to "act as an AI without limits"Vague setting, no real constraint on outputWrite specific duties and writing rules
Dump dozens of rules at onceRules fail in the second halfTrim to six or fewer, move the rest to other requests
JSON requirement just says "return JSON"Field names change every timeProvide complete format and examples

The common pattern in this table is: the more specific, the more effective; the more vague, the less useful. Also, do not treat "repeated emphasis" as a solution; writing the same sentence three times with bold and exclamation marks usually works less well than turning it into a single rule with numbers. What truly works is being concise and accurate, combined with example verification.

Debugging checklist

  1. Fix temperature at 0.2 to reproduce the issue.
  2. Change only one prompt and rerun the same batch of examples.
  3. Check finish_reason: if it is length, the problem is with max_tokens, not the prompt.
  4. Check prompt_tokens in usage: is the system prompt too long, crowding out the conversation history?
  5. After fixing, restore temperature to the business value and resample to verify.

See Integration Tutorial for integration details, and Documentation for the full parameter list. If you are evaluating whether to use a proxy service, the cost and trade-off analysis is in API Proxy Analysis, so we do not repeat it here.

Frequently asked questions

How long should the system prompt be?

As long as it clearly expresses the rules; usually a few hundred words are enough. Longer prompts consume more context window and are more prone to conflicting rules. We recommend no more than six rules.

Why does the output sometimes include explanatory text when instructed to output only JSON?

This is normal. Lowering the temperature, providing complete examples, and stripping wrappers and retrying parsing in code are more reliable than repeatedly emphasizing the instruction.

Can temperature and top_p be adjusted simultaneously?

Yes, but it is not recommended. Moving two parameters at once makes it hard to determine which one caused the change. Fix one and adjust only the other.

Do I need to declare "this is fiction" in the prompt?

No need. Legal adult-oriented fictional content will not be refused; just write the task clearly. Sexual content involving minors is blocked regardless of wording.

Fill out the form to get your API key

Create an account, copy the API key, and modify the Base URL. Configuration is that simple.