OpenAI /v1/chat/completions

POST /v1/chat/completions is cc-router’s OpenAI Chat Completions compatible entry point. Tools like Open WebUI, Cherry Studio, Cline, and LobeChat that only support “OpenAI-compatible” APIs can use every subscription you’ve aggregated just by pointing their Base URL at cc-router.

Applies to cc-router v5.0.0 and later. For a comparison of the three inbound endpoints, see API Overview.

Quick setup

SettingValue
Base URLhttp://127.0.0.1:23456/v1 (some tools want it without /v1 — follow the tool’s hint)
API KeyThe token from cc-router’s Settings page; if authentication is off, any non-empty value works
Model namemodel-fable / model-opus / model-sonnet / model-haiku, or aliases such as gpt-5.6 / gpt-5.5 / gpt-5.4 / gpt-5.4-mini

Most tools fetch the model list automatically via GET /v1/models — just pick one from it.

Open WebUI: Admin Panel → Settings → Connections → click + next to OpenAI API → set the URL to http://127.0.0.1:23456/v1 and the key to your token. If Open WebUI runs in Docker, replace 127.0.0.1 with host.docker.internal and switch cc-router’s Listen address to LAN.

Cherry Studio: Settings → Model Provider → Add → set the provider type to OpenAI → set the API host to http://127.0.0.1:23456 and the key to your token → Manage → fetch the model list.

OpenAI Python SDK:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:23456/v1", api_key="<cc-router token>")
resp = client.chat.completions.create(
    model="model-sonnet",
    messages=[{"role": "user", "content": "Explain cc-router in one sentence"}],
)
print(resp.choices[0].message.content)

Protocol position and translation flow

Client (Chat Completions)                     cc-router                              Upstream
─────────────────────────       ──────────────────────────────────────         ─────────────
POST /v1/chat/completions ─────►│ Chat → Anthropic Messages request    │
                                │   ↓                                  │
                                │ Dispatch pipeline (same as           │──► pick sub / rewrite model
                                │   /v1/messages)                      │◄── upstream response
                                │   ↓                                  │
                                │ Anthropic JSON / SSE → Chat format   │
◄────────────── HTTP response ──│                                      │

The dispatch pipeline is fully shared with /v1/messages: subscriptions, virtual models, quotas, session affinity, and automatic failover all apply. The upstream can be any outbound type (Anthropic / OpenAI / Gemini / OAuth) and the client can’t tell the difference. The Entry endpoint field in the request log details shows /v1/chat/completions.


Request

POST /v1/chat/completions
Content-Type: application/json
HeaderRequiredNotes
Content-Type: application/jsonYesThe body must be JSON
Authorization: Bearer <token> or x-api-keyPer auth settingsRequired when authentication is on
x-session-idNoSession identifier for session affinity; see below

Request fields

FieldBehavior
modelRequired. Resolved per the virtual model alias rules; missing returns 400
messagesRequired and must not be empty; see the translation rules in the next section
streamtrue uses SSE; defaults to false
max_completion_tokens / max_tokensTranslated to Anthropic max_tokens, with the former taking precedence; defaults to 4096 if neither is given
reasoning_effortTranslated to Anthropic thinking.budget_tokens per the table below; none means thinking stays off
temperature / top_pPassed through; not passed through when thinking is on (Anthropic requires temperature 1 with thinking, so the upstream default is used)
stopString or array, translated to stop_sequences
toolsOnly type: "function" is accepted; translated to Anthropic tool schema (parameters → input_schema)
tool_choiceauto / none / required (→ any) / a specific function (→ tool)
parallel_tool_callsfalse is translated to disable_parallel_tool_use: true
userTranslated to metadata.user_id, and also used as the session identifier for session affinity

reasoning_effort mapping

reasoning_effortthinking.budget_tokens
minimal1024
low2048
medium8192
high / xhigh / max16384
Any other value8192

When thinking is on and max_tokens is not greater than the budget, cc-router automatically raises max_tokens to “budget + 4096” to satisfy the upstream requirement that max_tokens exceed budget_tokens.

Forced tool calls take precedence over reasoning_effort: when tool_choice is required or names a specific function, cc-router keeps the tool choice and drops thinking. A forced tool call is a functional requirement while reasoning effort is only a preference, so the former wins when they conflict.

messages translation rules

Chat messageTranslated to Anthropic
role: system / developerMerged in order into the top-level system
role: user, texttext block
role: user, image_urlimage block; supports data:image/...;base64,... and http(s) URLs
role: assistant, contenttext block
role: assistant, tool_callstool_use block (arguments must be valid JSON)
role: assistant, reasoning_contentDropped. The client has no valid signature, so sending it back upstream would be rejected
role: tooltool_result block (must have tool_call_id); consecutive tool messages are merged into a single user message

Adjacent messages with the same role are merged automatically to satisfy Anthropic’s role-alternation requirement.

Unsupported or ignored fields

FieldHandling
Legacy functions / function_callReturns 400; use tools / tool_choice instead
Audio and file content partsReturns 400
n > 1, logprobs, logit_bias, seed, presence_penalty, frequency_penaltyIgnored
response_format (including JSON Schema enforcement)Ignored, no error
stream_optionsIgnored; streaming responses always end with a usage frame

Non-streaming request example

curl http://127.0.0.1:23456/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer <token>' \
  -d '{
    "model": "model-sonnet",
    "messages": [
      { "role": "system", "content": "You are concise." },
      { "role": "user", "content": "Explain cc-router in one sentence" }
    ]
  }' | jq

Response (non-streaming)

  • 200 OK, Content-Type: application/json, a standard chat.completion object
  • id is derived from the upstream message id: msg_abc → chatcmpl-abc
  • model echoes the model name the client sent (neither the virtual model name nor the real model name)
  • Thinking content goes in message.reasoning_content, following DeepSeek’s convention, so mainstream clients can display it collapsed
{
  "id": "chatcmpl-xxx",
  "object": "chat.completion",
  "created": 1790000000,
  "model": "model-sonnet",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "...",
        "reasoning_content": "..."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 42,
    "completion_tokens": 128,
    "total_tokens": 170,
    "prompt_tokens_details": { "cached_tokens": 0 }
  }
}

finish_reason mapping

Anthropic stop_reasonfinish_reason
end_turn / stop_sequencestop
max_tokenslength
tool_usetool_calls
refusalcontent_filter

Whenever the response contains tool calls, finish_reason is tool_calls (except for max_tokens / refusal), even if the upstream or relay reports end_turn. Clients can safely rely on finish_reason == "tool_calls" to decide whether to run tools.

usage mapping

Chat fieldComputed as
prompt_tokensinput_tokens + cache_creation_input_tokens + cache_read_input_tokens
completion_tokensoutput_tokens
total_tokensSum of the two
prompt_tokens_details.cached_tokenscache_read_input_tokens

Response (streaming SSE)

  • 200 OK, Content-Type: text/event-stream
  • Every frame is data: {chat.completion.chunk} with no event: line, and the stream ends with data: [DONE]

Frame sequence:

OrderContentSource Anthropic event
1delta: {"role": "assistant", "content": ""}message_start
…delta: {"content": "..."}text_delta
…delta: {"reasoning_content": "..."}thinking_delta
…delta: {"tool_calls": [{index, id, type, function: {name, arguments: ""}}]}Start of a tool_use block
…delta: {"tool_calls": [{index, function: {arguments: "..."}}]}input_json_delta
Third to lastdelta: {} + finish_reasonmessage_delta
Second to lastchoices: [] + usagemessage_stop
Lastdata: [DONE]

Streaming request example

curl -N http://127.0.0.1:23456/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer <token>' \
  -d '{
    "model": "model-sonnet",
    "stream": true,
    "messages": [{ "role": "user", "content": "ping" }]
  }'

Automatic retry on first-frame errors: same as /v1/messages — when the upstream returns 200 but the very first event is an error (e.g. quota exhaustion disguised as success), cc-router switches to the next subscription and retries, invisibly to the client.

Mid-stream errors and disconnect safety net: if an upstream error arrives mid-stream, cc-router emits one error frame plus data: [DONE]; if the upstream connection drops unexpectedly, it fills in finish_reason, usage, and data: [DONE] so the client never hangs waiting.


Session affinity

When the dispatch mode is Session affinity, cc-router identifies the session using the following priority, and pins each session to the same subscription:

  1. The user field in the request body (Open WebUI and others fill this in per user)
  2. The x-session-id request header
  3. A hash of the first user message’s content

If you’re writing your own client, passing a stable user or x-session-id per conversation noticeably improves the upstream prompt cache hit rate.


Error responses

Errors produced by /v1/chat/completions use the OpenAI shape:

{
  "error": {
    "message": "...",
    "type": "invalid_request_error",
    "code": "invalid_request_error",
    "param": null
  }
}

type is classified per OpenAI convention, while code keeps the original Anthropic error type to aid debugging:

Anthropic error type (code)OpenAI type
invalid_request_error / authentication_error / permission_errorinvalid_request_error
rate_limit_error / overloaded_errorrate_limit_error
Othersserver_error
StatusTrigger
400Body is not valid JSON, model is missing, messages is empty, or translation failed (legacy functions, unsupported content parts, tool_calls whose arguments isn’t valid JSON, etc.)
401Authentication failed. Note that the error body is Anthropic-shaped here, because authentication runs before the endpoint handler
500cc-router internal error, or failure to read / parse the upstream response
503The virtual model has no bound subscriptions, or every subscription is temporarily unavailable
Upstream statusWhen every subscription fails, the last upstream’s status code is passed through, with the error body translated to OpenAI shape