API documentation

OpenAI-compatible interface to 35 models from 7 upstream providers.

5-minute quickstart

Get from zero to a successful API call in five minutes. No SDK install required.

Before you copy any code: get the Base URL right

Every API request goes to this base — the trailing /v1 is required, not optional:

  • Base URL: https://api.houjiayan.com/v1
  • Chat: POST https://api.houjiayan.com/v1/chat/completions
  • Model list: GET https://api.houjiayan.com/v1/models

OpenAI-compatible clients disagree on this one field: the official SDKs, Cherry Studio and OpenClaw need the full https://api.houjiayan.com/v1; NextChat takes the bare domain and appends /v1 itself. Symptom of a wrong value: HTTP 200 with an HTML page instead of JSON — the most common integration failure reported by our integrators. Per-client values are spelled out in Client setup recipes below.

1. Create an account

Go to api.houjiayan.com/register. You'll need an email address; we send a verification link. New accounts receive a ¥6.6 (approx. $1.00) sign-up bonus — no fixed expiry; it stays valid as long as your account is active.

2. Generate an API key

In the dashboard, click "Add a new token". Give it a name (e.g. "dev-laptop") and copy the key — it's shown only once.

3. Make your first call

Run the curl below. You should see a JSON response with a 'Hello!' message in the assistant role.

curl https://api.houjiayan.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-flash",
    "messages": [
      {"role": "user", "content": "Hello"}
    ]
  }'

Prefer not to pick a model yourself? See Smart routing — set "model": "auto" and let the gateway choose.

Base URL

https://api.houjiayan.com/v1

Authentication

Pass your API key as a Bearer token in the Authorization header. Never expose your key in client-side code or public repositories.

Authorization: Bearer sk-hjy-...

Client setup recipes

Copy-paste configuration for popular OpenAI-compatible clients. Only the Base URL field differs per client — the exact value is noted for each.

Python (openai SDK)

base_url takes the full URL including /v1.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.houjiayan.com/v1"   # ← include /v1
)

Node.js (openai SDK)

baseURL takes the full URL including /v1.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "YOUR_API_KEY",
  baseURL: "https://api.houjiayan.com/v1"   // ← include /v1
});

LangChain (ChatOpenAI)

Pass base_url with the full URL including /v1 (older LangChain versions use the openai_api_base parameter).

from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="kimi-k2.6",
    api_key="YOUR_API_KEY",
    base_url="https://api.houjiayan.com/v1"   # older LangChain: openai_api_base=...
)

print(llm.invoke("Hello").content)

Cherry Studio

Add an OpenAI-compatible provider. The API host field takes the full URL including /v1.

Provider type : OpenAI (OpenAI-compatible)
API host    : https://api.houjiayan.com/v1   ← include /v1
API key     : sk-hjy-...
Models      : add by ID — e.g. kimi-k2.6, deepseek-flash

NextChat / ChatGPT-Next-Web

The custom endpoint (BASE_URL) takes the bare domain WITHOUT /v1 — NextChat appends /v1/chat/completions itself. Entering .../v1 here double-prefixes the path and requests fail.

Custom endpoint (BASE_URL) : https://api.houjiayan.com   ← bare domain, NO /v1
API key                  : sk-hjy-...
Custom models            : +kimi-k2.6,+deepseek-flash

OpenClaw

baseUrl takes the full URL including /v1 — OpenClaw does not append it. A domain-only value hits the gateway's web console and returns HTTP 200 with an HTML page instead of JSON; if you see that symptom, check this field first.

baseUrl : https://api.houjiayan.com/v1   ← include /v1 (not auto-appended)
apiKey  : sk-hjy-...
model   : kimi-k2.6   (or any ID from the models table)

cURL

No base URL field — call the full endpoint paths directly.

curl https://api.houjiayan.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-flash",
    "messages": [
      {"role": "user", "content": "Hello"}
    ]
  }'

Anthropic Messages format

Alongside the OpenAI-compatible format, the gateway natively speaks Anthropic Messages — the two formats coexist, so use whichever your tooling expects. Same models, same balance. Tools built on the Anthropic SDK, like Claude Code, plug in directly.

Same base URL as the OpenAI format — only the path and headers differ:

POST https://api.houjiayan.com/v1/messages

cURL

curl https://api.houjiayan.com/v1/messages \
  -H "x-api-key: YOUR_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-flash",
    "max_tokens": 1024,
    "messages": [
      {"role": "user", "content": "Hello"}
    ]
  }'

Python (anthropic SDK)

from anthropic import Anthropic

client = Anthropic(
    api_key="YOUR_API_KEY",
    base_url="https://api.houjiayan.com"
)

message = client.messages.create(
    model="deepseek-flash",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}]
)
print(message.content[0].text)

Claude Code

Claude Code works with any Anthropic-compatible endpoint. Point it at the gateway with two environment variables, then run claude as usual:

export ANTHROPIC_BASE_URL=https://api.houjiayan.com
export ANTHROPIC_AUTH_TOKEN=sk-hjy-...   # your platform API key

Available models

All models support the OpenAI Chat Completions schema. Pricing is per million tokens (1M tokens), set in CNY (¥) and shown here in USD ($, weekly reference rate). Input price is for prompt tokens; output price is for generated tokens. The four Kimi models are also in the smart routing pool — POST to /v1/auto/chat/completions with "model": "auto" and the gateway may pick a Kimi model automatically when it leads the quality band.

Model Provider Context Input Output Best for
deepseek-v4-pro DeepSeek 256K $1.37 $4.10 Flagship reasoning, complex analysis, research; peak pricing — 50% off off-peak (outside 9:00-12:00 & 14:00-18:00 Beijing weekdays)
deepseek-flash DeepSeek 1M $0.29 $1.17 Fast reasoning, chatbots, content generation; peak pricing — 50% off off-peak (outside 9:00-12:00 & 14:00-18:00 Beijing weekdays); cache hits billed at 1/50 of input
qwen3.7-max Alibaba Cloud 256K $1.82 $5.45 Flagship Qwen, translation, summarization
qwen3.7-plus Alibaba Cloud 128K $0.29 $1.18 Code generation, technical writing
qwen3.7-max-us Alibaba Cloud 256K $2.84 $8.52 US-region route of the Qwen flagship — same capability, lower price
qwen3.7-plus-us Alibaba Cloud 128K $0.45 $1.82 US-region route of qwen3.7-plus — everyday workloads at a lower price
qwen3.6-max-preview Alibaba Cloud 256K $1.36 $8.18 Next-gen Qwen flagship preview, translation, summarization. No authoritative benchmark yet — manual selection only, excluded from smart routing
qwen3.6-plus Alibaba Cloud 128K $0.30 $1.82 Next-gen Qwen mid-tier, code generation, technical writing. No authoritative benchmark yet — manual selection only, excluded from smart routing
qwen3.5-ocr Alibaba Cloud 32K $0.07 $0.29 Document OCR, image text extraction. No authoritative benchmark yet — manual selection only, excluded from smart routing
qwen3.8-flash Alibaba Cloud 262K (1M ext.) $0.11 $0.41 High-value Qwen flash, 125B-A6B, everyday chat and generation
qwen3.8-max Alibaba Cloud 1M $1.81 $5.42 Latest Qwen 2.4T-parameter MoE flagship — coding and office productivity, production-grade delivery
glm-5.3 Zhipu 1M $1.21 $4.24 Zhipu flagship, text — served via Zhipu official direct channel
glm-5.3-flash Zhipu 1M $0.11 $0.42 Zhipu multimodal (image/video/file), 320B-A18B
glm-5.2 Zhipu 1M $1.21 $4.24 Flagship for long-horizon tasks, lossless 1M context, engineering-grade coding
glm-5.2-fast-preview Zhipu 1M $2.42 $8.48 High-speed GLM-5.2 variant — capability parity, 1.5–2× output TPS
glm-5.2-us Zhipu 1M $1.59 $5.00 US-region deployment of glm-5.2 (Model Studio Virginia)
glm-5.1 Zhipu 200K $1.21 $4.24 Complex code and long-horizon tasks — autonomous runs up to 8 hours
MiniMax-M3 MiniMax 1M $0.64 $2.55 Flagship multimodal, vision + text — priced at the 1M context tier; ≤512K context tier: $0.32/$1.27
MiniMax-M2.7 MiniMax 256K $0.32 $1.27 Tool calling, agent systems, multi-step planning
mimo-v2.6-flash Xiaomi 1M $0.14 $0.28 Omni-modal efficient reasoning (text/image/video/audio in), low-cost high-frequency workloads; cache hits $0.0028/1M
mimo-v2.5-asr Xiaomi 8K $3.38 $3.38 Speech recognition (audio in / text out), Chinese-English + dialects, open-sourced
mimo-v2.6-pro Xiaomi 1M $0.435 $0.87 Xiaomi's most powerful omni-modal flagship — long-horizon complex workflows; cache hits $0.0036/1M
mimo-v2.6-pro-ultraspeed Xiaomi 1M $4.35 $8.70 V2.6-Pro flagship performance at up to 20× output speed — real-time production; cache hits $0.036/1M
mimo-v2.5-tts Xiaomi 8K $0.00 $0.00 Text-to-speech, premium preset voices, style instructions and audio tags
mimo-v2.5-tts-voiceclone Xiaomi — $0.00 $0.00 Voice cloning from a few seconds of reference audio
mimo-v2.5-tts-voicedesign Xiaomi — $0.00 $0.00 Voice design — create a new voice from a text description
doubao-seed-2-1-pro-260628 ByteDance 256K $0.91 $4.55 Doubao Seed 2.1 flagship Pro — deep thinking + multimodal, coding/agent
doubao-seed-2-1-turbo-260628 ByteDance 256K $0.45 $2.27 Lower-cost, low-latency Seed 2.1 variant — deep thinking + multimodal
doubao-seed-evolving ByteDance 1M $0.91 $4.55 Latest Seed coding & agent model — unified ID, auto-upgrades weekly
doubao-seed-character-260628 ByteDance 128K $0.12~0.18 $0.30~0.91 Role-play/persona model — tiered by context (≤32K / 32K–128K)
doubao-seed-translation-250915 ByteDance 4K $0.18 $0.55 7B multilingual translation, 28 languages — Responses API only
kimi-k3 Moonshot 1M $2.95 $14.75 Flagship all-round, top-tier coding, long-context agents
kimi-k2.7-code Moonshot 256K $0.98 $4.09 Dedicated coding, repo-level edits, debugging
kimi-k2.7-code-highspeed Moonshot 256K $1.97 $8.18 High-speed coding, low-latency dev loops
kimi-k2.6 Moonshot 256K $0.98 $4.09 High-value general purpose, everyday chat + writing

Pin a model, or let the gateway route

Two ways to set the model field. (1) Pin an exact ID from the table above (e.g. "model": "kimi-k2.6") — the request goes straight to that model via /v1/chat/completions. (2) Smart routing: POST to /v1/auto/chat/completions with "model": "auto" and an optional "preference": "quality" (default — best result), "balanced" (cost-aware inside the quality band) or "cost" (cheapest, quality not guaranteed). Routed responses include a routing_decision block so you can audit every choice.

curl https://api.houjiayan.com/v1/auto/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "preference": "balanced",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Prefer not to pick a model yourself? See Smart routing — set "model": "auto" and let the gateway choose.

Code examples

cURL

curl https://api.houjiayan.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-flash",
    "messages": [
      {"role": "user", "content": "Hello"}
    ]
  }'

Python (openai SDK)

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.houjiayan.com/v1"
)

response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[{"role": "user", "content": "Hello"}]
)
print(response.choices[0].message.content)

Node.js (openai SDK)

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "YOUR_API_KEY",
  baseURL: "https://api.houjiayan.com/v1"
});

const response = await client.chat.completions.create({
  model: "deepseek-flash",
  messages: [{ role: "user", content: "Hello" }]
});
console.log(response.choices[0].message.content);

Streaming response (SSE)

Set stream: true to receive tokens as they are generated. Each chunk is a server-sent event with a delta field.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.houjiayan.com/v1"
)

stream = client.chat.completions.create(
    model="deepseek-flash",
    messages=[{"role": "user", "content": "Write a haiku about APIs"}],
    stream=True
)
for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

Function calling

Most models support OpenAI-compatible tool calling. Define your tools in the tools array; the model returns structured tool_calls that you execute and feed back as tool role messages.

from openai import OpenAI
import json

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.houjiayan.com/v1")

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather in a city",
        "parameters": {
            "type": "object",
            "properties": {
                "city": {"type": "string", "description": "City name"}
            },
            "required": ["city"]
        }
    }
}]

response = client.chat.completions.create(
    model="MiniMax-M2.7",
    messages=[{"role": "user", "content": "What's the weather in Hohhot?"}],
    tools=tools
)
tool_call = response.choices[0].message.tool_calls[0]
print(f"Call: {tool_call.function.name}({tool_call.function.arguments})")

Error codes

  • 401 Invalid or missing API key. Check the Authorization header and that the key has not been revoked.
  • 402 Insufficient balance. Top up via the dashboard, or request a refund if the 24h window applies.
  • 403 API key does not have access to this model, or account is suspended for Terms violation.
  • 404 Model name is unknown. Use /v1/models to list all available models.
  • 413 Request payload too large. Reduce prompt size or use a model with larger context window.
  • 429 Rate limit hit. The gateway auto-retries on a different upstream where possible. If persistent, see Rate limits below.
  • 500/502/503/504 Upstream provider error. The gateway returns the error from the source. Try a different model, or retry with exponential backoff.
  • 529 All upstream providers are overloaded. The gateway could not complete the request. Retry after 30 seconds.

Rate limits

Per-user limits: 60 requests/minute, 100K tokens/minute (input + output combined). The gateway auto-failover ensures that if one upstream returns 429, traffic shifts to another provider where available. If you need higher limits, email support@houjiayan.com with your use case.

Concurrent requests

Up to 5 concurrent in-flight requests per API key. Exceeding this returns 429. For batch workloads, use the /v1/batch endpoint (coming Q4 2026).

Best practices

  • Always set explicit max_tokens to avoid runaway costs (especially on reasoning models).
  • Use stream: true for user-facing chat UIs to reduce perceived latency.
  • Cache identical prompts for 5-60 seconds at the application layer to reduce token spend on retries.
  • Implement exponential backoff with jitter for 429 and 5xx errors.
  • Never put your API key in client-side code. Proxy through your backend.
  • For long documents, prefer MiniMax-M3 or deepseek-flash (1M context) over chunking.

Need help?

Email support@houjiayan.com or open a ticket in the dashboard. We respond within 24 hours on business days.