API Overview: Inbound and Outbound
cc-router sits between your tools and the model providers:
- Inbound: how your tools connect to cc-router. cc-router exposes endpoints for three protocols — point each tool at whichever one it speaks.
- Outbound: how cc-router connects to providers. Each subscription is configured with the protocol its provider supports, and cc-router sends correctly formatted requests.
The two sides are independent and can be combined freely. For example, Codex can come in through the OpenAI Responses endpoint and be answered by DeepSeek’s Anthropic endpoint.
Claude Code OpenCode OpenClaw pi ... Codex ... Open WebUI / Cherry Studio ...
| | | | | |
------------------------------------- | |
| | |
Anthropic OpenAI OpenAI
Messages API Responses API Chat Completions API
(/v1/messages) (/v1/responses) (/v1/chat/completions)
| | |
-------------------------------------------------------
| Inbound · virtual models
|
cc-router
(local 127.0.0.1:23456)
|
| Outbound · real models
-----------------------------------------------------------------------------
| | | | | | |
DeepSeek GLM Kimi Anthropic OpenAI Gemini ......
API Coding Coding Messages Responses & API
Plan Plan API Completions
Internally, everything becomes Anthropic Messages
cc-router has a single dispatch pipeline inside, and it speaks the Anthropic Messages protocol:
- A request arrives at any inbound endpoint. Non-Anthropic inbound requests are first translated into Anthropic Messages.
- The pipeline resolves
modelto a virtual model (model-fable/model-opus/model-sonnet/model-haiku) and picks a subscription based on that slot’s bound subscription list and dispatch mode (sequential / round-robin / session affinity). - The request is sent using that subscription’s outbound protocol. Non-Anthropic outbound protocols get another translation step.
- The response is translated back along the same path and returned in the client’s inbound protocol format.
So all three inbound endpoints share the same subscriptions, virtual models, quotas, and session affinity; rate limiting, automatic retries on failure, and subscription failover apply to every endpoint. The Entry endpoint field in the request log details shows which endpoint each request came in through.
Inbound: how your tools connect to cc-router
| Inbound endpoint (click for full reference) | Typical clients | Base URL | Protocol translation |
|---|---|---|---|
Anthropic MessagesPOST /v1/messages | Claude Code, Claude Desktop, OpenCode, OpenClaw, pi, Kimi Code CLI | http://127.0.0.1:23456 | None, passed through as-is |
OpenAI ResponsesPOST /v1/responses | Codex CLI, Codex Desktop | http://127.0.0.1:23456/v1 | Responses ⇄ Messages |
OpenAI Chat CompletionsPOST /v1/chat/completions | Open WebUI, Cherry Studio, Cline, LobeChat | http://127.0.0.1:23456/v1 | Chat ⇄ Messages |
Besides the three inbound endpoints, there are two shared endpoints:
| Endpoint | Purpose | Auth |
|---|---|---|
GET /v1/models | Virtual model list; fields are compatible with both Anthropic and OpenAI SDKs | Not required |
GET /health | Liveness probe; returns plain-text ok | Not required |
See GET /v1/models for details.
Quick comparison of the three inbound endpoints
/v1/messages | /v1/responses | /v1/chat/completions | |
|---|---|---|---|
| Minimum version | v3.0.0 | v3.0.0 | v5.0.0 |
| Recommended model names | model-opus etc. | gpt-5.5 etc. | Either works |
| Reasoning effort | thinking / output_config.effort passed through as-is | reasoning.effort → thinking budget | reasoning_effort → thinking budget |
| Returned reasoning content | thinking blocks (signed) | reasoning items (signed, can be sent back) | reasoning_content field (dropped if sent back) |
| Image input | Supported | Not supported | Supported (data: base64 and http(s) URLs) |
| Tool calling | Supported | Supported (function type) | Supported (tools / tool_choice; legacy functions not supported) |
| Stream terminator | message_stop | response.completed (no [DONE]) | data: [DONE] |
| Streaming usage | In message_start / message_delta | In response.completed | Always one usage frame at the end |
| Error body style | Anthropic | OpenAI | OpenAI |
| Session affinity key | x-claude-code-session-id → metadata.user_id → first user message | prompt_cache_key → session_id header | user → x-session-id header → first user message |
Which endpoint should I use? If your tool supports Anthropic Messages, prefer it: zero translation, and thinking,
cache_control, images, and tool calls all keep their native semantics. Use the other two endpoints only when your tool speaks nothing but OpenAI protocols.
Shared conventions
Listening and ports
| Setting | Default | Notes |
|---|---|---|
| Listen address | 127.0.0.1 | Becomes 0.0.0.0 after switching Settings → Proxy service → Listen address to LAN |
| HTTP port | 23456 | Auto-increments by 1 if taken, up to 100 attempts |
| HTTPS port | 23457 | Enabled when Listen protocol is HTTPS only or HTTP + HTTPS; also auto-increments |
| Request body limit | 32 MB | Requests with many base64 images may exceed it and get 413; you can raise it in Settings |
Changes to the port, protocol, listen address, or request body limit all require restarting cc-router to take effect.
Authentication
- Token authentication is on by default (Settings → Authentication & CORS). When on, every inbound endpoint reads the token from
x-api-key: <token>orAuthorization: Bearer <token>(x-api-keywins), and it must exactly match the token in Settings. /v1/models,/health, and allOPTIONSpreflight requests don’t require authentication.- The
401body for auth failures is always Anthropic-shaped, regardless of the inbound endpoint, because authentication runs before the endpoint handler. - This token only grants access to cc-router and has nothing to do with upstream providers’ API keys. cc-router swaps in the upstream key per subscription when forwarding.
HTTPS
The HTTPS port uses a certificate issued by cc-router’s local self-signed CA, so clients must trust that CA first. You only need it for clients that accept HTTPS exclusively (such as Claude Desktop). See Claude Desktop Integration (macOS) and the Windows edition for setup.
CORS
On by default: Access-Control-Allow-Origin: *, methods GET / POST / OPTIONS, preflight returns 204 directly. 401 responses carry CORS headers too, so a browser can read the error body.
Virtual models and aliases
Every inbound endpoint shares the same model-name resolution rules, and model names can be mixed across endpoints:
| Virtual model | Accepted aliases |
|---|---|
model-fable | claude-fable*, gpt-5.6, gpt-*-sol |
model-opus | claude-opus*, gpt-5.5, gpt-*-terra |
model-sonnet | claude-sonnet*, gpt-5.4, gpt-*-luna |
model-haiku | claude-haiku*, gpt-*-mini |
- Any of the names above can take an
anthropic/oropenai/prefix (LiteLLM style) with the same effect. claude-opus*is a prefix match:claude-opus-4-7andclaude-opus-4-7-20260101both resolve tomodel-opus.gpt-*-solmatches the tier by--separated segments:gpt-5.6-solandgpt-6-sol-20261201both resolve tomodel-fable;gpt-5.4-miniresolves tomodel-haiku.- Model names that match no rule go to fallback:
modelis passed through as-is to the fallback slot’s subscriptions.
Outbound: how cc-router connects to providers
Outbound connections fall into four categories by upstream protocol. Built-in provider presets and custom endpoints take the same path; presets simply come with the address, auth method, and model list pre-filled. The full list of built-in presets is whatever the in-app Add subscription page shows.
| Outbound type | Upstream protocol | Built-in presets (excerpt) | Custom endpoints | Protocol translation |
|---|---|---|---|---|
| Anthropic Messages compatible | /v1/messages | Anthropic (official), DeepSeek, Zhipu GLM, Kimi, MiniMax, Xiaomi MiMo, Alibaba Cloud Bailian, Volcengine Ark, Tencent Cloud, Baidu Qianfan, StepFun, ModelScope, UCloud, Fireworks, OpenRouter, xAI, Ollama, and more | Any Anthropic Messages compatible endpoint | None, passed through as-is |
| OpenAI compatible | /v1/responses, /v1/chat/completions | OpenAI official API | one-api / new-api relays, Groq, Together, local vLLM / llama.cpp, etc. | Messages → OpenAI |
| Gemini compatible | generateContent, /v1beta/interactions | Google AI Studio, Gemini Interactions API | Any Gemini compatible endpoint | Messages → Gemini |
| Subscription accounts (OAuth) | Provider-specific private protocols | Codex (ChatGPT Plus/Pro), Kiro (AWS) | N/A | Messages → private protocol |
Key behaviors of each outbound type:
- Anthropic Messages compatible: the main path. No protocol translation — thinking,
output_config.effort,cache_control, images, and tool calls all keep their native semantics. When a provider offers a native Anthropic endpoint, configure this type first. - OpenAI compatible: thinking and OpenAI reasoning are mapped in both directions, and multi-turn reasoning context is fed back automatically;
reasoning_contentreturned by Chat Completions upstreams (DeepSeek R1, etc.) is converted into thinking blocks. Anything the translation layer can’t express (e.g.cache_control) is dropped. - Gemini compatible: thinking is mapped in both directions, and thought signatures are carried automatically across tool-call round-trips.
- Subscription accounts (OAuth): no API key; you sign in via the OAuth device-code flow. This is a gray area with a risk of account bans — don’t rely on it as your main path; it’s only suitable as a last-resort fallback.
Inbound × outbound combinations
Any inbound endpoint can pair with any outbound type; the difference is how many protocol translations happen in between:
| Inbound ↓ / Outbound → | Anthropic compatible | OpenAI compatible / Gemini compatible / OAuth |
|---|---|---|
/v1/messages | 0, full passthrough, highest fidelity | 1 (outbound translation) |
/v1/responses | 1 (inbound translation) | 2 |
/v1/chat/completions | 1 (inbound translation) | 2 |
The more translations, the more likely protocol-specific features get lost (for example, cache_control only works when the whole path is Anthropic). In practice, one or two translations make little difference for everyday chat and tool calling, but if you care about prompt cache hit rates or reasoning details, keep both inbound and outbound on the Anthropic protocol.