EN ▾
Get API Key

Wu Xianzhi APIMigration Guide

API Gateway Migration Guide: Switch from OpenAI, OpenRouter to an Uncensored API

If your project currently uses OpenAI, OpenRouter, or an API relay and you want to switch to an uncensored API that won't reject valid requests, you only need to change three configurations: base_url, API key, and model name. This article first clarifies the difference between a relay and a dedicated uncensored model, then provides a parameter comparison table, code for parallel switching using environment variables, a pre-launch checklist, and the most common pitfalls during migration.

Updated on

Key Points

  1. Gateways resell original models with unchanged content policies; only dedicated uncensored models solve refusal issues
  2. Migration requires only three changes: set base_url to https://api.wuxianzhiapi.com/v1,密钥,模型名 uncensored
  3. No support for embeddings, images, audio, or fine-tuning; continue using your original service for these
  4. Use environment variables for a gray release; if issues arise, changing one variable allows a rollback.

What is the difference between an API gateway and a dedicated uncensored model

Clarify the concepts first, or you might choose the wrong direction during migration. A common API relay essentially resells or aggregates usage quotas from major models, providing an OpenAI-compatible address so you can switch models using the same SDK. It solves the problem of access and billing, such as a unified entry point and billing, but the model itself remains the original one. The original provider's content policies remain unchanged: topics that are rejected will still be rejected, and changing the relay address does not change this.

A dedicated uncensored model is different. It is not forwarding another model; it is a separately provided model where valid adult content, fictional creation, and controversial topics are not rejected. Wu Xianzhi API provides only one model, with the model name uncensored. The endpoint is compatible with the OpenAI format, so migration costs are low, but it has clear boundaries: it handles text only, does not support images, audio, vectors, or fine-tuning; sexual content involving minors is blocked (returning 403) regardless of whether it is fictional.

So before migrating, ask yourself: is your problem an unstable endpoint or high prices, or does the model always reject your valid requests? If it is the latter, switching relays makes little sense; switching to a dedicated uncensored API is the right solution. Many teams keep both: general tasks continue using the original endpoint, while requests needing uncensored output are routed here. The following sections explain how to do this.

Three things to change when migrating from OpenAI or OpenRouter

Whether you use OpenAI, OpenRouter, or a relay, if your code uses an OpenAI-compatible SDK, change three things: base_url to https://api.wuxianzhiapi.com/v1; api_key to the key from /get-api-key/; model to uncensored. No other models are available; GET /v1/models shows only this one.

import os
from openai import OpenAI

messages = [{"role": "user", "content": "你好"}]

# 迁移前(示意):
# client = OpenAI(api_key=os.environ["OLD_API_KEY"], base_url="旧地址")
# resp = client.chat.completions.create(model="旧模型名", messages=messages)

# 迁移后:只动 base_url、api_key、model 这三处
client = OpenAI(
    base_url="https://api.wuxianzhiapi.com/v1",
    api_key=os.environ["WUXIANZHI_API_KEY"],
)
resp = client.chat.completions.create(model="uncensored", messages=messages, max_tokens=100)
print(resp.choices[0].message.content)

For Node.js, replace new OpenAI({...})'s baseURL and apiKey. If you use HTTP requests directly, change the address to https://api.wuxianzhiapi.com/v1/chat/completions and keep the header Authorization: Bearer <key>. See code examples for Python, Node.js, and cURL.

curl https://api.wuxianzhiapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WUXIANZHI_API_KEY" \
  -d '{"model":"uncensored","messages":[{"role":"user","content":"你好"}],"max_tokens":50}'

Parameter mapping: what works and what doesn't

The table below covers fields commonly encountered during migration. The rule is: core fields related to chat and OpenAI format work as usual; features related to "other models or modalities" are not supported here.

Original usageHow to handle here
model (e.g., various GPT models)Must change to uncensored
messages (system / user / assistant / tool)Format is identical; use directly
max_tokensDefault 2048, max 32,000; returns 400 if exceeded
stream: trueSupported; appends a usage block at the end
tools / tool_choiceSupported, OpenAI format
Context lengthTotal prompt + output: 100,000 tokens
Request body sizeMax 8 MB
Rate limit300 requests per minute per API key
Vector embeddingsNot supported
Image generation / vision, audio, videoNot supported; text only
Fine-tuningNot supported
Switch between multiple modelsOnly one model is available; there is no list to switch between.

For optional fields not listed, do not assume they behave as the original factory/provider way. Test them separately in a test environment to confirm expected behavior before going live. See API documentation for supported features.

What to do if you lack certain capabilities: alternatives for vectors, images, and audio

If your original project uses both chat and vector retrieval, do not try to switch everything. We only provide text chat, so embeddings-related code, such as knowledge base retrieval and semantic deduplication, must continue using your original vector service or switch to a self-deployed vector solution. Switching the chat part here while keeping the retrieval part unchanged is the easiest separation, as the two sides do not affect each other.

The same applies to images and audio. For example, if your product is text with images, text generation can go here while image generation continues using your original image endpoint; for text-to-speech, hand over to your existing speech service after text generation. Extracting the text generation step into a separate function means that no matter how you combine other capabilities later, the scope of changes will be minimal.

If your code uses multiple models, e.g., cheap for classification and expensive for creation, here only one model handles classification. Input unit price is $0.25 / 1M tokens, so cost is low. Set max_tokens small to ignore cost. See pricing page.

Parallel run: switch between two endpoints using environment variables

The biggest fear during migration is a "big bang" cut-over. A more stable approach is to create a thin wrapper layer in your code, using environment variables to decide which API to call. This allows you to route a small portion of traffic or specific features to the new API first. If issues arise, you can revert by changing a single variable. Since both APIs follow the OpenAI-compatible format, the wrapper is very simple to implement.

import os
from openai import OpenAI

PROVIDERS = {
    "old": {
        "base_url": os.environ.get("OLD_BASE_URL", ""),
        "api_key": os.environ.get("OLD_API_KEY", ""),
        "model": os.environ.get("OLD_MODEL", ""),
    },
    "wuxianzhi": {
        "base_url": "https://api.wuxianzhiapi.com/v1",
        "api_key": os.environ.get("WUXIANZHI_API_KEY", ""),
        "model": "uncensored",
    },
}

def get_client(name=None):
    name = name or os.environ.get("LLM_PROVIDER", "old")
    cfg = PROVIDERS[name]
    return OpenAI(base_url=cfg["base_url"], api_key=cfg["api_key"]), cfg["model"]

def chat(messages, provider=None, **kwargs):
    client, model = get_client(provider)
    return client.chat.completions.create(model=model, messages=messages, **kwargs)

# 通用任务走旧接口,需要无审查输出的请求显式指定新接口
resp = chat([{"role": "user", "content": "写一个黑色幽默的短故事"}],
            provider="wuxianzhi", max_tokens=800)
print(resp.choices[0].message.content)

Switch granularity has three levels: by environment (test first), by feature (only creation endpoints), or by user (gray-scale a portion of accounts). Keep the provider field in logs to troubleshoot differences. Store chat history as standard messages array for seamless continuation.

Migration checklist

Go through the following steps in order before going live to ensure nothing is missed:

  1. Register at /get-api-key/, get your key, and use $0.50 free trial credit (valid for 7 days) to verify. No top up required.
  2. Use curl /v1/models to confirm the key is valid and the network is reachable.
  3. Drive base_url, api_key, and model via environment variables so the key is not committed to the code repository.
  4. Search your code for hardcoded model names, max_tokens values, and embeddings calls.
  5. Check streaming code to handle the final choices empty usage block.
  6. Add exponential backoff retries for 429 and 503 errors, and add explicit error handling branches for 402 and 403 errors.
  7. Run a batch of regression samples using your real prompts, focusing on whether previously rejected prompts now output correctly.
  8. Start with a small percentage of traffic, compare usage and latency, and scale up only if everything looks good.
  9. Confirm that your product targets adult users and that the use case is legal; this is a prerequisite for using this API.

Common pitfalls during migration

Forgot to change the model name. Sending old GPT models or aggregator paths returns errors. Globally search for the model name and ensure you send uncensored.

max_tokens exceeded. Some projects set max_tokens to 32000+ for long outputs, but max here is 32,000 (returns 400). Prompt + max_tokens cannot exceed 100,000 tokens; lower the output limit accordingly for long inputs.

Streaming usage blocks. Before the stream ends, the server automatically appends a block containing usage, where choices is an empty array. If your parsing code directly accesses chunk.choices[0], it will throw an error at the final step. Some older code also manually passes stream_options to request usage data; this is not needed here.

Mistaking "uncensored" for "unbounded". Legal adult content, fiction, and controversial topics are not rejected, but sexual content involving minors is always blocked, including in fiction and roleplay, returning a 403 with content_blocked. The product must handle adult user access control itself.

Balance and trial expiration. The trial credit expires after 7 days. When the balance is exhausted, you will receive a 402 error with the code no_credit. Translate this error into a user-friendly message in your application instead of a generic "service error". For specific scenario designs, refer to the application scenarios article.

Frequently Asked Questions

Do I need to rewrite my original prompts after migration?

The format does not need to change; the messages structure is exactly the same. However, you can delete the jailbreak-style preamble that was previously written to bypass rejections. Clearly stating the role and task saves tokens and is more stable.

Can I keep both old and new APIs during migration?

Yes, and we recommend doing so. Use environment variables to determine the base_url, key, and model name. Route some features or users to the new API first. If issues arise, you can revert by changing a single variable.

What happens if I migrate embeddings calls from the old API?

There is no embeddings API here; requests will return 404. Continue using your original service for vector retrieval and only migrate the chat requests.

How do I know if the migration improved performance?

Run a regression comparison using a batch of real prompts that were previously rejected or rewritten. Record the rejection rate, response length, and token counts from the usage data. The samples should come from your own business data, not from generic online test sets.

Simply fill out the form to obtain your key

Create an account, copy your key, and modify the Base URL. Configuration is that simple.

Get API Key