A smart LLM gateway that plans with a strong model and executes with cheap ones — so you save tokens without losing quality.
Point any OpenAI-compatible tool (Codex, Cursor, Aider, Continue) at Switchboard and it decides, per request, how to get a good answer for the least money.
Модели — как работники: дорогая (Opus) = опытный спец, дешёвая = стажёр. Switchboard — это «прораб» посередине. Он сам решает, кого позвать на каждую задачу, а для сложного просит дорогую модель написать короткий план и отдаёт работу дешёвым. Ты ничего не переключаешь руками — оно само, вживую.
Most gateways (LiteLLM, RouteLLM) pick one model per request. Switchboard also supports plan → execute: a strong model writes a short plan, cheap models do the bulk of the work, then a final synthesis. That is where the real token savings on hard tasks come from.
Strategies are chosen by the model field of the request:
model value |
What it does | Status |
|---|---|---|
auto-cascade |
Try cheap first; escalate to strong only if answer shaky | ✅ MVP |
auto-classify |
Judge difficulty first, then pick the model in one pass | ✅ MVP |
auto-plan |
Strong plans, cheap executes, then synthesize | ✅ MVP |
| a concrete id | Pass straight through to that model | ✅ |
Runs with zero API keys using fake "mock" models, so you can see it work immediately.
bun install
bun startIn another terminal:
curl http://localhost:4000/v1/chat/completions -H "content-type: application/json" -d "{\"model\":\"auto-cascade\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello!\"}]}"The (non-streaming) JSON response includes a switchboard block showing the route, the cost, and how much was saved.
Copy .env.example to .env and add whichever key you have:
# Anthropic (cheap/strong auto-picks claude-haiku-4-5 / claude-opus-4-8)
ANTHROPIC_API_KEY=sk-ant-...
# or OpenAI (auto-picks gpt-4o-mini / gpt-4o)
OPENAI_API_KEY=sk-...
The cheap/strong pair is chosen automatically from your keys; override with
CHEAP_MODEL / STRONG_MODEL to pick any models from src/registry.ts
(e.g. claude-haiku-4-5, claude-sonnet-5, claude-opus-4-8, claude-fable-5).
Route the easy majority of your work to free local models (via Ollama) and save an expensive cloud model (or a rate-limited subscription) for the hard parts. Nothing leaves your machine, and no API tokens or subscription limits are spent on routine work.
ollama pull qwen2.5-coder:7b # capable local "strong" model (planner)
ollama pull qwen3.5:2b # fast local "cheap" model (executor)
bun run start:local # cheap=qwen3.5:2b, strong=qwen2.5-coder:7b, all $0Every request is now answered locally for $0. Point a tool at it (below) and your day-to-day coding assistant runs free — you only reach for the paid/subscription model when the task is genuinely hard.
Small local models (7B) are great for chat and routine coding, but heavy agentic tools may need a larger model. Pull
qwen2.5-coder:14bfor more capability (slower on ≤8 GB VRAM).
export OPENAI_BASE_URL="http://localhost:4000/v1"
export OPENAI_API_KEY="anything" # Switchboard uses your provider keys from .envThen run Codex / Cursor / Aider as usual — every request now flows through Switchboard.
For Anthropic-format clients (Claude Code and other Claude SDKs):
export ANTHROPIC_BASE_URL="http://localhost:4000"
export ANTHROPIC_API_KEY="anything"Switchboard serves both APIs: POST /v1/chat/completions (OpenAI) and
POST /v1/messages (Anthropic) — same routing engine behind each.
Every request is logged to a local SQLite file (switchboard.db). Ask the gateway
for cumulative totals:
curl http://localhost:4000/statsYou get requests, tokens, % of tokens handled by cheap models, dollars spent vs
the strong-only baseline, dollars saved and saved %, plus a per-strategy breakdown.
Add ?since=<unix-seconds> to window it (e.g. this month). This is the raw material
for a "you saved $X this month" dashboard or subscription view.
See the savings and the routing quality for yourself:
bun run eval # default strategy: auto-classify
EVAL_STRATEGY=auto-cascade bun run evalIt runs a labelled prompt set through the router and reports routing accuracy (did easy prompts go cheap and hard prompts go strong?) and % cost saved vs running the strong model on everything:
--- summary (8 cases) ---
routing accuracy: 100% (8/8 sent to the right tier)
SAVED: 33%
Point it at your real models (via .env) for real dollar figures. This is the
"save X% without losing quality" claim, measured — and it makes the savings-vs-accuracy
trade-off between strategies explicit (cascade saves more but leans on the confidence
signal; classify routes more precisely).
Add EVAL_JUDGE=1 and a capable judge model to prove the router's cheaper answers
are as good as the strong model's:
EVAL_JUDGE=1 EVAL_JUDGE_MODEL=gpt-4o bun run evalFor each case it fetches a reference answer from the strong model and asks the judge
whether the router's answer is ACCEPTABLE or WORSE, then reports quality retained %
next to the savings — the full "saved X% without losing quality" proof. (A mock judge
reports n/a; point it at a real model for real scores.)
Everything works from env vars and sensible defaults, but you can drop a
switchboard.config.json (copy switchboard.config.example.json) to set the models,
the classifier threshold, or register custom models with pricing:
{
"cheapModel": "gpt-4o-mini",
"strongModel": "gpt-4o",
"classifyThreshold": 0.4,
"models": {
"my-custom-model": { "provider": "openai", "tier": "strong", "inputPerM": 5, "outputPerM": 15 }
}
}Precedence is env var > config file > built-in default. Point elsewhere with
SWITCHBOARD_CONFIG=/path/to/config.json.
- OpenAI-compatible
/v1/chat/completions - Cascade strategy + cost/savings reporting
-
auto-plan(plan → execute) strategy -
auto-classify(judge difficulty first) strategy - Cumulative stats (SQLite) +
/statsendpoint (saved $, tokens, per strategy) - Streaming (SSE) responses — OpenAI-compatible chunks (so Cursor/Codex/SDKs work)
- Real, verified pricing table (Anthropic + OpenAI)
- Eval harness (
bun run eval) — routing accuracy + % cost saved - LLM-judge quality scoring in eval (
EVAL_JUDGE=1) — prove quality is unchanged - Config file (
switchboard.config.json) — models, classifier threshold, custom pricing - True incremental (token-by-token) streaming
- Anthropic-native
/v1/messagesendpoint (Claude Code / Claude SDK clients) - Provider fallback (same-tier, cross-provider) on errors
Early MVP. Prices in src/registry.ts are real published rates (Anthropic verified
2026-06-24, OpenAI 2026) — still confirm current pricing for your own account/region.
MIT licensed.