EN ▾
Get API key

Wu Xianzhi APICode Examples

Uncensored AI API call examples: Python, Node.js, and cURL full code

This is a developer-friendly manual for the uncensored AI API. The endpoint is compatible with OpenAI Chat Completions, so your familiar openai SDK works by changing two lines of configuration. This article covers the most common tasks at once: basic requests, streaming, function calling, error retry, cost control with max_tokens, and maintaining context in multi-turn conversations. All code is copy-paste ready; keys are read from environment variables.

Updated on

Key points

  1. Change base_url to https://api.wuxianzhiapi.com/v1,模型名写 uncensored; the official openai SDK needs no other changes
  2. The last streaming block is usage data with an empty choices array; check for emptiness before reading
  3. Retry 429 and 503 with exponential backoff; retrying 400/401/402/403 is pointless
  4. The API is stateless; manage history yourself and use max_tokens and truncation to control costs

API basics and environment variables

Memorize these fixed parameters; all examples use them. Base URL is https://api.wuxianzhiapi.com/v1, model name is fixed as uncensored, auth is header Authorization: Bearer <key>. Only two endpoints: POST /v1/chat/completions for chat, GET /v1/models to check availability. Request/response matches OpenAI Chat Completions, so official openai SDK only needs base_url and key changes; business code stays mostly unchanged.

The API key is displayed immediately after registration at /get-api-key/. Sign in with email and password. New accounts receive $0.50 in free trial credit, valid for 7 days, with no card required. All examples in this article read the key from the environment variable WUXIANZHI_API_KEY. Do not commit the key to your code repository or put it in frontend pages. A few limits to clarify upfront: context window of 100,000 tokens (prompt and output combined), request body max 8 MB, and 300 requests per minute per API key.

export WUXIANZHI_API_KEY="把你的密钥放这里"

# 确认连通性,应返回包含 uncensored 的模型列表
curl https://api.wuxianzhiapi.com/v1/models \
  -H "Authorization: Bearer $WUXIANZHI_API_KEY"

If you get 401, the key is wrong or the env var isn't set. Debug this first. See API docs for full parameter details.

cURL: minimal working request

Regardless of your final language, we suggest running cURL first to get it working. This isolates network, key, and request format issues from your business code. Below is a standard request with max_tokens. The response JSON contains the reply in choices[0].message.content and token usage in usage. Billing is calculated based on these two numbers.

curl https://api.wuxianzhiapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WUXIANZHI_API_KEY" \
  -d '{
    "model": "uncensored",
    "messages": [
      {"role": "system", "content": "你是一个直来直去的写作助手。"},
      {"role": "user", "content": "用三句话描述一场暴雨前的小镇。"}
    ],
    "max_tokens": 300
  }'

Note: Chinese in JSON needs no manual escaping; just declare UTF-8 compatible JSON in headers; most terminals allow direct paste. If debugging in Windows PowerShell where quote handling is tricky, save request body to body.json and send with -d @body.json.

Python and Node.js full calls

Python uses official openai package (v1+), install with pip install openai. Pass base_url and api_key to client; usage is identical to OpenAI. Save script as chat.py to run.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.wuxianzhiapi.com/v1",
    api_key=os.environ["WUXIANZHI_API_KEY"],
)

resp = client.chat.completions.create(
    model="uncensored",
    messages=[
        {"role": "system", "content": "你是一个直来直去的写作助手。"},
        {"role": "user", "content": "用三句话描述一场暴雨前的小镇。"},
    ],
    max_tokens=300,
)

print(resp.choices[0].message.content)
print("输入 tokens:", resp.usage.prompt_tokens, "输出 tokens:", resp.usage.completion_tokens)

Use the openai npm package (v4 or later) with Node.js. Run npm install openai. The code below uses top-level await, so save the file as chat.mjs or set "type": "module" in package.json.

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.wuxianzhiapi.com/v1",
  apiKey: process.env.WUXIANZHI_API_KEY,
});

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [
    { role: "system", content: "你是一个直来直去的写作助手。" },
    { role: "user", content: "用三句话描述一场暴雨前的小镇。" },
  ],
  max_tokens: 300,
});

console.log(resp.choices[0].message.content);
console.log("用量:", resp.usage);

Code structure is identical; only syntax differs. If you already use OpenAI, replace client init lines and change model to uncensored. See migration guide for full migration steps.

How to read streaming (SSE)

Use streaming for long text or chat UIs so users see output incrementally. Set stream: true. The server pushes SSE chunks (one line each: data: {...}), ending with data: [DONE]. The official SDK parses this; you just iterate.

Pitfall: stream ends with a usage block where choices is empty. No extra param needed, but don't access chunk.choices[0] directly; check for emptiness first to avoid index errors.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.wuxianzhiapi.com/v1",
    api_key=os.environ["WUXIANZHI_API_KEY"],
)

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "写一段 200 字左右的悬疑小说开头。"}],
    max_tokens=600,
    stream=True,
)

usage = None
for chunk in stream:
    if chunk.usage:            # 最后一块:用量统计
        usage = chunk.usage
    if not chunk.choices:      # 用量块没有 choices
        continue
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)

print()
if usage:
    print("输入", usage.prompt_tokens, "输出", usage.completion_tokens)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.wuxianzhiapi.com/v1",
  apiKey: process.env.WUXIANZHI_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "写一段 200 字左右的悬疑小说开头。" }],
  max_tokens: 600,
  stream: true,
});

let usage = null;
for await (const chunk of stream) {
  if (chunk.usage) usage = chunk.usage;
  const delta = chunk.choices?.[0]?.delta?.content;
  if (delta) process.stdout.write(delta);
}
console.log("\n用量:", usage);

To see raw SSE data, use cURL with -N to disable buffering. This prints chunks immediately, useful for debugging if a proxy or gateway swallows the stream.

curl -N https://api.wuxianzhiapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WUXIANZHI_API_KEY" \
  -d '{"model":"uncensored","stream":true,"max_tokens":100,
       "messages":[{"role":"user","content":"数到五。"}]}'

If forwarding streams via Nginx or similar, disable response buffering for that path; otherwise, the frontend sees a single burst instead of incremental output.

Function calling: tools and tool results

Function calling uses OpenAI format: declare functions with tools (name, description, JSON Schema). Model returns name and JSON args in message.tool_calls. Execute function, return result as role: "tool" message, then request again for final answer.

tool_choice defaults to "auto". Force function with {"type": "function", "function": {"name": "get_weather"}}; disable with "none". Example: first request gets tool_calls, execute locally, second request sends tool results.

import json, os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.wuxianzhiapi.com/v1",
    api_key=os.environ["WUXIANZHI_API_KEY"],
)

def get_weather(city: str) -> dict:
    # 这里用假数据代替真实的天气接口
    return {"city": city, "temp_c": 18, "condition": "多云"}

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "查询指定城市的当前天气",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string", "description": "城市名"}},
            "required": ["city"],
        },
    },
}]

messages = [{"role": "user", "content": "杭州现在天气怎么样?"}]

first = client.chat.completions.create(
    model="uncensored", messages=messages, tools=tools, tool_choice="auto", max_tokens=500
)
msg = first.choices[0].message

if msg.tool_calls:
    messages.append(msg)  # 必须把带 tool_calls 的 assistant 消息原样放回历史
    for call in msg.tool_calls:
        args = json.loads(call.function.arguments)
        result = get_weather(**args)
        messages.append({
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps(result, ensure_ascii=False),
        })
    final = client.chat.completions.create(
        model="uncensored", messages=messages, tools=tools, max_tokens=500
    )
    print(final.choices[0].message.content)
else:
    print(msg.content)

Three errors: 1. Forget to append assistant tool_calls to history before tool message. 2. Mismatched tool_call_id. 3. Args are strings; json.loads them and handle errors; don't inject directly into SQL/commands. Node.js flow is identical.

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.wuxianzhiapi.com/v1",
  apiKey: process.env.WUXIANZHI_API_KEY,
});

const getWeather = (city) => ({ city, temp_c: 18, condition: "多云" });

const tools = [{
  type: "function",
  function: {
    name: "get_weather",
    description: "查询指定城市的当前天气",
    parameters: {
      type: "object",
      properties: { city: { type: "string", description: "城市名" } },
      required: ["city"],
    },
  },
}];

const messages = [{ role: "user", content: "杭州现在天气怎么样?" }];

const first = await client.chat.completions.create({
  model: "uncensored", messages, tools, tool_choice: "auto", max_tokens: 500,
});
const msg = first.choices[0].message;

if (msg.tool_calls?.length) {
  messages.push(msg);
  for (const call of msg.tool_calls) {
    const args = JSON.parse(call.function.arguments);
    messages.push({
      role: "tool",
      tool_call_id: call.id,
      content: JSON.stringify(getWeather(args.city)),
    });
  }
  const final = await client.chat.completions.create({
    model: "uncensored", messages, tools, max_tokens: 500,
  });
  console.log(final.choices[0].message.content);
} else {
  console.log(msg.content);
}

Error handling and retries: 429 and 503 backoff

Error responses are all unified JSON: {"error":{"code":...,"message":...}}. In your code, you only need to retry two types: 429 (exceeding the 300 requests per minute rate limit) and 503 (upstream_busy, model temporarily busy, try again in a few seconds). Network-level connection timeouts are also worth retrying. Retrying other errors is futile: 400 means the request itself is problematic (e.g., prompt plus max_tokens exceeds 100k), 401 means the API key is invalid, 402 no_credit means balance is exhausted or trial expired, 403 content_blocked means content is blocked, and resending a hundred times yields the same result.

Use exponential backoff with jitter: wait ~1s for the 1st attempt, ~2s for the 2nd, ~4s for the 3rd. Set a limit and max count to prevent concurrent tasks from retrying simultaneously, which worsens rate limiting. The official SDK has max_retries, which retries 429 and 5xx errors by default; for simple scenarios, just increase it. For logging, circuit breaking, or custom wait times, write your own loop.

import os, random, time
import openai
from openai import OpenAI

# max_retries=0:关闭 SDK 自带重试,完全由下面的函数控制
client = OpenAI(
    base_url="https://api.wuxianzhiapi.com/v1",
    api_key=os.environ["WUXIANZHI_API_KEY"],
    max_retries=0,
    timeout=60,
)

def chat_with_retry(messages, max_attempts=5, **kwargs):
    for attempt in range(max_attempts):
        try:
            return client.chat.completions.create(
                model="uncensored", messages=messages, **kwargs
            )
        except (openai.RateLimitError, openai.InternalServerError,
                openai.APIConnectionError, openai.APITimeoutError) as e:
            if attempt == max_attempts - 1:
                raise
            wait = min(30, 2 ** attempt) + random.uniform(0, 1)
            print(f"{type(e).__name__},{wait:.1f} 秒后重试(第 {attempt + 1} 次)")
            time.sleep(wait)
        except openai.APIStatusError as e:
            # 400 / 401 / 402 / 403 / 404:重试无效,直接交给上层处理
            print("不可重试:", e.status_code, e.response.text)
            raise

resp = chat_with_retry([{"role": "user", "content": "你好"}], max_tokens=100)
print(resp.choices[0].message.content)

Node.js: increase maxRetries directly; SDK handles backoff for 429/5xx. Check error.status to distinguish errors.

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.wuxianzhiapi.com/v1",
  apiKey: process.env.WUXIANZHI_API_KEY,
  maxRetries: 5,
  timeout: 60_000,
});

try {
  const resp = await client.chat.completions.create({
    model: "uncensored",
    messages: [{ role: "user", content: "你好" }],
    max_tokens: 100,
  });
  console.log(resp.choices[0].message.content);
} catch (err) {
  if (err instanceof OpenAI.APIError) {
    if (err.status === 402) console.error("余额不足,请充值后再试");
    else if (err.status === 403) console.error("内容被拦截:", err.message);
    else console.error("请求失败:", err.status, err.message);
  } else {
    throw err;
  }
}

Note: if stream breaks, keep received content. Retries regenerate from start and re-bill. For long text, use segmented requests instead of one large request.

Control cost and length with max_tokens

Billing is straightforward: input $0.25 per million tokens, output $1.00 per million tokens. Prepaid credit, no monthly fee, balance never expires. Output costs 4 times the input, so the real savings are in the output. max_tokens defaults to 2048, with a single request max of 32,000. If your scenario only needs a sentence or two of reply, set it explicitly to 200 or 300 to prevent the model from rambling and to cap the cost per request.

If truncated, finish_reason is "length", meaning max_tokens limit reached, not model completion. Check this field to decide whether to continue writing.

Multi-turn chat: manage context yourself

API is stateless; server doesn't remember previous requests. To continue chat, resend full history in order: system, then alternating user/assistant. Each turn adds input tokens, increasing cumulative cost.

The total context is limited to 100,000 tokens (including this output), so long conversations must be truncated. The simplest approach is to keep the system message and the most recent few turns; for a more complex approach, you can summarize earlier content in a single request into a short summary and place it in the system message. The class below encapsulates the logic for saving history and truncating by count, ready to be dropped into your chat service.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.wuxianzhiapi.com/v1",
    api_key=os.environ["WUXIANZHI_API_KEY"],
)

class Chat:
    def __init__(self, system: str, keep_last: int = 20):
        self.system = {"role": "system", "content": system}
        self.history = []          # 只存 user / assistant 消息
        self.keep_last = keep_last

    def say(self, text: str, max_tokens: int = 500) -> str:
        self.history.append({"role": "user", "content": text})
        recent = self.history[-self.keep_last:]
        resp = client.chat.completions.create(
            model="uncensored",
            messages=[self.system] + recent,
            max_tokens=max_tokens,
        )
        reply = resp.choices[0].message.content
        self.history.append({"role": "assistant", "content": reply})
        return reply

bot = Chat("你是一位说话简短的旅行顾问。")
print(bot.say("我想去云南玩五天,有什么建议?"))
print(bot.say("刚才说的第二个地方,适合带老人吗?"))  # 能接上上一轮

Truncating by message count is sufficient but imprecise, since message lengths vary greatly. If you need strict control, use the usage.prompt_tokens from the response as a reference for actual usage: proactively compress history once it nears 50,000. If you need to build a long-term companion product, see the context design examples in use cases.

Frequently asked questions

Why does the last streaming data block have no content?

That is the automatically appended usage statistics block, with an empty choices array and token counts in usage. When reading, first check if choices is empty, then take the delta. No need to pass extra parameters to enable it.

Should both 429 and 503 be retried? How long should you wait?

Both are worth retrying. 429 indicates exceeding the 300 requests per minute limit, while 503 with upstream_busy means the model is temporarily busy. Exponential backoff with random jitter is recommended, starting at 1 second, with a maximum retry count to avoid infinite retries.

What if the model doesn't return tool_calls during function calling?

This means the model determined no call was needed; the message.content is the final answer. If a call is mandatory, set tool_choice to a specific function and check whether the function description and parameter schema are clearly defined.

Will multi-turn conversations get more expensive?

Yes. The API is stateless, so history must be resent each time, and input tokens accumulate with each turn. Keep only the most recent few turns, or compress early content into a summary, while limiting output with max_tokens.

Fill out the form to get your key

Create an account, copy the key, and update the Base URL. Configuration is that simple.

Get API key