API Overview: Inbound and Outbound

cc-router sits between your tools and the model providers:

  • Inbound: how your tools connect to cc-router. cc-router exposes endpoints for three protocols — point each tool at whichever one it speaks.
  • Outbound: how cc-router connects to providers. Each subscription is configured with the protocol its provider supports, and cc-router sends correctly formatted requests.

The two sides are independent and can be combined freely. For example, Codex can come in through the OpenAI Responses endpoint and be answered by DeepSeek’s Anthropic endpoint.

 Claude Code    OpenCode    OpenClaw   pi ...   Codex ...      Open WebUI / Cherry Studio ...
      |             |           |         |         |                         |
      -------------------------------------         |                         |
                        |                           |                         |
                    Anthropic                    OpenAI                    OpenAI
                  Messages API                Responses API         Chat Completions API
                 (/v1/messages)              (/v1/responses)       (/v1/chat/completions)
                        |                           |                         |
                        -------------------------------------------------------
                                                  |  Inbound · virtual models
                                                  |
                                              cc-router
                                        (local 127.0.0.1:23456)
                                                  |
                                                  |  Outbound · real models
           -----------------------------------------------------------------------------
           |            |            |            |            |            |          |
       DeepSeek        GLM         Kimi       Anthropic     OpenAI       Gemini     ......
          API        Coding       Coding      Messages    Responses &      API
                      Plan         Plan          API      Completions

Internally, everything becomes Anthropic Messages

cc-router has a single dispatch pipeline inside, and it speaks the Anthropic Messages protocol:

  1. A request arrives at any inbound endpoint. Non-Anthropic inbound requests are first translated into Anthropic Messages.
  2. The pipeline resolves model to a virtual model (model-fable / model-opus / model-sonnet / model-haiku) and picks a subscription based on that slot’s bound subscription list and dispatch mode (sequential / round-robin / session affinity).
  3. The request is sent using that subscription’s outbound protocol. Non-Anthropic outbound protocols get another translation step.
  4. The response is translated back along the same path and returned in the client’s inbound protocol format.

So all three inbound endpoints share the same subscriptions, virtual models, quotas, and session affinity; rate limiting, automatic retries on failure, and subscription failover apply to every endpoint. The Entry endpoint field in the request log details shows which endpoint each request came in through.

Inbound: how your tools connect to cc-router

Inbound endpoint (click for full reference)Typical clientsBase URLProtocol translation
Anthropic Messages
POST /v1/messages
Claude Code, Claude Desktop, OpenCode, OpenClaw, pi, Kimi Code CLIhttp://127.0.0.1:23456None, passed through as-is
OpenAI Responses
POST /v1/responses
Codex CLI, Codex Desktophttp://127.0.0.1:23456/v1Responses ⇄ Messages
OpenAI Chat Completions
POST /v1/chat/completions
Open WebUI, Cherry Studio, Cline, LobeChathttp://127.0.0.1:23456/v1Chat ⇄ Messages

Besides the three inbound endpoints, there are two shared endpoints:

EndpointPurposeAuth
GET /v1/modelsVirtual model list; fields are compatible with both Anthropic and OpenAI SDKsNot required
GET /healthLiveness probe; returns plain-text okNot required

See GET /v1/models for details.

Quick comparison of the three inbound endpoints

/v1/messages/v1/responses/v1/chat/completions
Minimum versionv3.0.0v3.0.0v5.0.0
Recommended model namesmodel-opus etc.gpt-5.5 etc.Either works
Reasoning effortthinking / output_config.effort passed through as-isreasoning.effort → thinking budgetreasoning_effort → thinking budget
Returned reasoning contentthinking blocks (signed)reasoning items (signed, can be sent back)reasoning_content field (dropped if sent back)
Image inputSupportedNot supportedSupported (data: base64 and http(s) URLs)
Tool callingSupportedSupported (function type)Supported (tools / tool_choice; legacy functions not supported)
Stream terminatormessage_stopresponse.completed (no [DONE])data: [DONE]
Streaming usageIn message_start / message_deltaIn response.completedAlways one usage frame at the end
Error body styleAnthropicOpenAIOpenAI
Session affinity keyx-claude-code-session-id → metadata.user_id → first user messageprompt_cache_key → session_id headeruser → x-session-id header → first user message

Which endpoint should I use? If your tool supports Anthropic Messages, prefer it: zero translation, and thinking, cache_control, images, and tool calls all keep their native semantics. Use the other two endpoints only when your tool speaks nothing but OpenAI protocols.

Shared conventions

Listening and ports

SettingDefaultNotes
Listen address127.0.0.1Becomes 0.0.0.0 after switching Settings → Proxy service → Listen address to LAN
HTTP port23456Auto-increments by 1 if taken, up to 100 attempts
HTTPS port23457Enabled when Listen protocol is HTTPS only or HTTP + HTTPS; also auto-increments
Request body limit32 MBRequests with many base64 images may exceed it and get 413; you can raise it in Settings

Changes to the port, protocol, listen address, or request body limit all require restarting cc-router to take effect.

Authentication

  • Token authentication is on by default (Settings → Authentication & CORS). When on, every inbound endpoint reads the token from x-api-key: <token> or Authorization: Bearer <token> (x-api-key wins), and it must exactly match the token in Settings.
  • /v1/models, /health, and all OPTIONS preflight requests don’t require authentication.
  • The 401 body for auth failures is always Anthropic-shaped, regardless of the inbound endpoint, because authentication runs before the endpoint handler.
  • This token only grants access to cc-router and has nothing to do with upstream providers’ API keys. cc-router swaps in the upstream key per subscription when forwarding.

HTTPS

The HTTPS port uses a certificate issued by cc-router’s local self-signed CA, so clients must trust that CA first. You only need it for clients that accept HTTPS exclusively (such as Claude Desktop). See Claude Desktop Integration (macOS) and the Windows edition for setup.

CORS

On by default: Access-Control-Allow-Origin: *, methods GET / POST / OPTIONS, preflight returns 204 directly. 401 responses carry CORS headers too, so a browser can read the error body.

Virtual models and aliases

Every inbound endpoint shares the same model-name resolution rules, and model names can be mixed across endpoints:

Virtual modelAccepted aliases
model-fableclaude-fable*, gpt-5.6, gpt-*-sol
model-opusclaude-opus*, gpt-5.5, gpt-*-terra
model-sonnetclaude-sonnet*, gpt-5.4, gpt-*-luna
model-haikuclaude-haiku*, gpt-*-mini
  • Any of the names above can take an anthropic/ or openai/ prefix (LiteLLM style) with the same effect.
  • claude-opus* is a prefix match: claude-opus-4-7 and claude-opus-4-7-20260101 both resolve to model-opus.
  • gpt-*-sol matches the tier by --separated segments: gpt-5.6-sol and gpt-6-sol-20261201 both resolve to model-fable; gpt-5.4-mini resolves to model-haiku.
  • Model names that match no rule go to fallback: model is passed through as-is to the fallback slot’s subscriptions.

Outbound: how cc-router connects to providers

Outbound connections fall into four categories by upstream protocol. Built-in provider presets and custom endpoints take the same path; presets simply come with the address, auth method, and model list pre-filled. The full list of built-in presets is whatever the in-app Add subscription page shows.

Outbound typeUpstream protocolBuilt-in presets (excerpt)Custom endpointsProtocol translation
Anthropic Messages compatible/v1/messagesAnthropic (official), DeepSeek, Zhipu GLM, Kimi, MiniMax, Xiaomi MiMo, Alibaba Cloud Bailian, Volcengine Ark, Tencent Cloud, Baidu Qianfan, StepFun, ModelScope, UCloud, Fireworks, OpenRouter, xAI, Ollama, and moreAny Anthropic Messages compatible endpointNone, passed through as-is
OpenAI compatible/v1/responses, /v1/chat/completionsOpenAI official APIone-api / new-api relays, Groq, Together, local vLLM / llama.cpp, etc.Messages → OpenAI
Gemini compatiblegenerateContent, /v1beta/interactionsGoogle AI Studio, Gemini Interactions APIAny Gemini compatible endpointMessages → Gemini
Subscription accounts (OAuth)Provider-specific private protocolsCodex (ChatGPT Plus/Pro), Kiro (AWS)N/AMessages → private protocol

Key behaviors of each outbound type:

  • Anthropic Messages compatible: the main path. No protocol translation — thinking, output_config.effort, cache_control, images, and tool calls all keep their native semantics. When a provider offers a native Anthropic endpoint, configure this type first.
  • OpenAI compatible: thinking and OpenAI reasoning are mapped in both directions, and multi-turn reasoning context is fed back automatically; reasoning_content returned by Chat Completions upstreams (DeepSeek R1, etc.) is converted into thinking blocks. Anything the translation layer can’t express (e.g. cache_control) is dropped.
  • Gemini compatible: thinking is mapped in both directions, and thought signatures are carried automatically across tool-call round-trips.
  • Subscription accounts (OAuth): no API key; you sign in via the OAuth device-code flow. This is a gray area with a risk of account bans — don’t rely on it as your main path; it’s only suitable as a last-resort fallback.

Inbound × outbound combinations

Any inbound endpoint can pair with any outbound type; the difference is how many protocol translations happen in between:

Inbound ↓ / Outbound →Anthropic compatibleOpenAI compatible / Gemini compatible / OAuth
/v1/messages0, full passthrough, highest fidelity1 (outbound translation)
/v1/responses1 (inbound translation)2
/v1/chat/completions1 (inbound translation)2

The more translations, the more likely protocol-specific features get lost (for example, cache_control only works when the whole path is Anthropic). In practice, one or two translations make little difference for everyday chat and tool calling, but if you care about prompt cache hit rates or reasoning details, keep both inbound and outbound on the Anthropic protocol.