OpenAI /v1/chat/completions
POST /v1/chat/completions is cc-router’s OpenAI Chat Completions compatible entry point. Tools like Open WebUI, Cherry Studio, Cline, and LobeChat that only support “OpenAI-compatible” APIs can use every subscription you’ve aggregated just by pointing their Base URL at cc-router.
Applies to cc-router v5.0.0 and later. For a comparison of the three inbound endpoints, see API Overview.
Quick setup
| Setting | Value |
|---|---|
| Base URL | http://127.0.0.1:23456/v1 (some tools want it without /v1 — follow the tool’s hint) |
| API Key | The token from cc-router’s Settings page; if authentication is off, any non-empty value works |
| Model name | model-fable / model-opus / model-sonnet / model-haiku, or aliases such as gpt-5.6 / gpt-5.5 / gpt-5.4 / gpt-5.4-mini |
Most tools fetch the model list automatically via GET /v1/models — just pick one from it.
Open WebUI: Admin Panel → Settings → Connections → click + next to OpenAI API → set the URL to http://127.0.0.1:23456/v1 and the key to your token. If Open WebUI runs in Docker, replace 127.0.0.1 with host.docker.internal and switch cc-router’s Listen address to LAN.
Cherry Studio: Settings → Model Provider → Add → set the provider type to OpenAI → set the API host to http://127.0.0.1:23456 and the key to your token → Manage → fetch the model list.
OpenAI Python SDK:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:23456/v1", api_key="<cc-router token>")
resp = client.chat.completions.create(
model="model-sonnet",
messages=[{"role": "user", "content": "Explain cc-router in one sentence"}],
)
print(resp.choices[0].message.content)
Protocol position and translation flow
Client (Chat Completions) cc-router Upstream
───────────────────────── ────────────────────────────────────── ─────────────
POST /v1/chat/completions ─────►│ Chat → Anthropic Messages request │
│ ↓ │
│ Dispatch pipeline (same as │──► pick sub / rewrite model
│ /v1/messages) │◄── upstream response
│ ↓ │
│ Anthropic JSON / SSE → Chat format │
◄────────────── HTTP response ──│ │
The dispatch pipeline is fully shared with /v1/messages: subscriptions, virtual models, quotas, session affinity, and automatic failover all apply. The upstream can be any outbound type (Anthropic / OpenAI / Gemini / OAuth) and the client can’t tell the difference. The Entry endpoint field in the request log details shows /v1/chat/completions.
Request
POST /v1/chat/completions
Content-Type: application/json
| Header | Required | Notes |
|---|---|---|
Content-Type: application/json | Yes | The body must be JSON |
Authorization: Bearer <token> or x-api-key | Per auth settings | Required when authentication is on |
x-session-id | No | Session identifier for session affinity; see below |
Request fields
| Field | Behavior |
|---|---|
model | Required. Resolved per the virtual model alias rules; missing returns 400 |
messages | Required and must not be empty; see the translation rules in the next section |
stream | true uses SSE; defaults to false |
max_completion_tokens / max_tokens | Translated to Anthropic max_tokens, with the former taking precedence; defaults to 4096 if neither is given |
reasoning_effort | Translated to Anthropic thinking.budget_tokens per the table below; none means thinking stays off |
temperature / top_p | Passed through; not passed through when thinking is on (Anthropic requires temperature 1 with thinking, so the upstream default is used) |
stop | String or array, translated to stop_sequences |
tools | Only type: "function" is accepted; translated to Anthropic tool schema (parameters → input_schema) |
tool_choice | auto / none / required (→ any) / a specific function (→ tool) |
parallel_tool_calls | false is translated to disable_parallel_tool_use: true |
user | Translated to metadata.user_id, and also used as the session identifier for session affinity |
reasoning_effort mapping
reasoning_effort | thinking.budget_tokens |
|---|---|
minimal | 1024 |
low | 2048 |
medium | 8192 |
high / xhigh / max | 16384 |
| Any other value | 8192 |
When thinking is on and max_tokens is not greater than the budget, cc-router automatically raises max_tokens to “budget + 4096” to satisfy the upstream requirement that max_tokens exceed budget_tokens.
Forced tool calls take precedence over
reasoning_effort: whentool_choiceisrequiredor names a specific function, cc-router keeps the tool choice and drops thinking. A forced tool call is a functional requirement while reasoning effort is only a preference, so the former wins when they conflict.
messages translation rules
| Chat message | Translated to Anthropic |
|---|---|
role: system / developer | Merged in order into the top-level system |
role: user, text | text block |
role: user, image_url | image block; supports data:image/...;base64,... and http(s) URLs |
role: assistant, content | text block |
role: assistant, tool_calls | tool_use block (arguments must be valid JSON) |
role: assistant, reasoning_content | Dropped. The client has no valid signature, so sending it back upstream would be rejected |
role: tool | tool_result block (must have tool_call_id); consecutive tool messages are merged into a single user message |
Adjacent messages with the same role are merged automatically to satisfy Anthropic’s role-alternation requirement.
Unsupported or ignored fields
| Field | Handling |
|---|---|
Legacy functions / function_call | Returns 400; use tools / tool_choice instead |
| Audio and file content parts | Returns 400 |
n > 1, logprobs, logit_bias, seed, presence_penalty, frequency_penalty | Ignored |
response_format (including JSON Schema enforcement) | Ignored, no error |
stream_options | Ignored; streaming responses always end with a usage frame |
Non-streaming request example
curl http://127.0.0.1:23456/v1/chat/completions \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <token>' \
-d '{
"model": "model-sonnet",
"messages": [
{ "role": "system", "content": "You are concise." },
{ "role": "user", "content": "Explain cc-router in one sentence" }
]
}' | jq
Response (non-streaming)
200 OK,Content-Type: application/json, a standardchat.completionobjectidis derived from the upstream message id:msg_abc→chatcmpl-abcmodelechoes the model name the client sent (neither the virtual model name nor the real model name)- Thinking content goes in
message.reasoning_content, following DeepSeek’s convention, so mainstream clients can display it collapsed
{
"id": "chatcmpl-xxx",
"object": "chat.completion",
"created": 1790000000,
"model": "model-sonnet",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "...",
"reasoning_content": "..."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 42,
"completion_tokens": 128,
"total_tokens": 170,
"prompt_tokens_details": { "cached_tokens": 0 }
}
}
finish_reason mapping
Anthropic stop_reason | finish_reason |
|---|---|
end_turn / stop_sequence | stop |
max_tokens | length |
tool_use | tool_calls |
refusal | content_filter |
Whenever the response contains tool calls, finish_reason is tool_calls (except for max_tokens / refusal), even if the upstream or relay reports end_turn. Clients can safely rely on finish_reason == "tool_calls" to decide whether to run tools.
usage mapping
| Chat field | Computed as |
|---|---|
prompt_tokens | input_tokens + cache_creation_input_tokens + cache_read_input_tokens |
completion_tokens | output_tokens |
total_tokens | Sum of the two |
prompt_tokens_details.cached_tokens | cache_read_input_tokens |
Response (streaming SSE)
200 OK,Content-Type: text/event-stream- Every frame is
data: {chat.completion.chunk}with noevent:line, and the stream ends withdata: [DONE]
Frame sequence:
| Order | Content | Source Anthropic event |
|---|---|---|
| 1 | delta: {"role": "assistant", "content": ""} | message_start |
| … | delta: {"content": "..."} | text_delta |
| … | delta: {"reasoning_content": "..."} | thinking_delta |
| … | delta: {"tool_calls": [{index, id, type, function: {name, arguments: ""}}]} | Start of a tool_use block |
| … | delta: {"tool_calls": [{index, function: {arguments: "..."}}]} | input_json_delta |
| Third to last | delta: {} + finish_reason | message_delta |
| Second to last | choices: [] + usage | message_stop |
| Last | data: [DONE] |
Streaming request example
curl -N http://127.0.0.1:23456/v1/chat/completions \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <token>' \
-d '{
"model": "model-sonnet",
"stream": true,
"messages": [{ "role": "user", "content": "ping" }]
}'
Automatic retry on first-frame errors: same as
/v1/messages— when the upstream returns200but the very first event is an error (e.g. quota exhaustion disguised as success), cc-router switches to the next subscription and retries, invisibly to the client.Mid-stream errors and disconnect safety net: if an upstream error arrives mid-stream, cc-router emits one error frame plus
data: [DONE]; if the upstream connection drops unexpectedly, it fills infinish_reason, usage, anddata: [DONE]so the client never hangs waiting.
Session affinity
When the dispatch mode is Session affinity, cc-router identifies the session using the following priority, and pins each session to the same subscription:
- The
userfield in the request body (Open WebUI and others fill this in per user) - The
x-session-idrequest header - A hash of the first
usermessage’s content
If you’re writing your own client, passing a stable user or x-session-id per conversation noticeably improves the upstream prompt cache hit rate.
Error responses
Errors produced by /v1/chat/completions use the OpenAI shape:
{
"error": {
"message": "...",
"type": "invalid_request_error",
"code": "invalid_request_error",
"param": null
}
}
type is classified per OpenAI convention, while code keeps the original Anthropic error type to aid debugging:
Anthropic error type (code) | OpenAI type |
|---|---|
invalid_request_error / authentication_error / permission_error | invalid_request_error |
rate_limit_error / overloaded_error | rate_limit_error |
| Others | server_error |
| Status | Trigger |
|---|---|
400 | Body is not valid JSON, model is missing, messages is empty, or translation failed (legacy functions, unsupported content parts, tool_calls whose arguments isn’t valid JSON, etc.) |
401 | Authentication failed. Note that the error body is Anthropic-shaped here, because authentication runs before the endpoint handler |
500 | cc-router internal error, or failure to read / parse the upstream response |
503 | The virtual model has no bound subscriptions, or every subscription is temporarily unavailable |
| Upstream status | When every subscription fails, the last upstream’s status code is passed through, with the error body translated to OpenAI shape |