LLM tooling · for Claude Code

Make Claude Code punch above its quota.

brainmux builds tools that route your work to the right model — so Claude Code spends your Opus quota on architecture and review, not busywork.

brainmux · open source
# one brand, a family of Claude Code tools
▸ llmproxy live
▸ more soon

What is it

brainmux is LLM tooling for Claude Code. Its first tool, llmproxy, routes Claude Code to cheap OpenRouter models through local LiteLLM proxies and delegates the grunt work to them — one OpenRouter key reaches thousands of models, the brains run pay-as-you-go and never touch your Anthropic quota, and Opus stays the orchestrator.

Products

One tool today. A family in the making.

llmproxy

Live · v0.1

Run Claude Code on cheap LLM brains and delegate the grunt work — one OpenRouter key, thousands of models, your Opus quota untouched.

Explore llmproxy ↓

more brainmux tools

Soon

Same idea, new surfaces. Built in the open under brainmux.

llmproxy · our first tool

Run Claude Code on cheap brains.

Your Opus quota is the bottleneck. llmproxy routes Claude Code to cheap OpenRouter models and delegates the grunt work to them — Opus stays the orchestrator; cheap brains do the volume.

install · Claude Code
>/plugin marketplace add brainmuxhq/brainmux
✓ added marketplace: brainmux
>/plugin install llmproxy@brainmux
✓ installed llmproxy
one key

Thousands of models

A single OpenRouter key reaches DeepSeek, Qwen, GLM, GPT, Gemini and hundreds more — no per-provider setup.

separate meter

Never touches your quota

The brains run pay-as-you-go on OpenRouter. Your Anthropic subscription quota stays untouched.

orchestrator

Opus stays in charge

Cheap brains do the volume — bulk edits, detection sweeps. Opus reviews, decides, and fixes.

How it works

One config. One proxy per brain. Routed by port.

A single brains.yaml (zod-validated) is the source of truth. bmux generates a Docker Compose stack — one LiteLLM proxy per brain, isolated by port — and points Claude Code at the brain you choose.

brains.yaml — your brains (SSOT)
bmux generate
compose · litellm configs · init.sql
bmux up
chat :4567 · deep :4568 · coder :4569
Claude Code, on your cheap brain

What you get

A CLI for the whole loop.

bmux cli

Manage the stack

init · up / down · health · config add-brain / set-model · test. Declarative — edit, regenerate, restart.

delegate

Offload grunt work

bmux delegate coder "…" hands a bounded task to a cheap brain headless. Opus verifies the result — no rubber-stamp.

bmux models

Pick from the live catalog

Browse the live OpenRouter catalog by price, context and use-case — never a guessed, stale model slug.

observability

Spend & logs, per brain

Each brain ships the LiteLLM UI for spend, request logs and parameter tuning. Nothing to rebuild.

Quickstart

Four commands to a running brain.

Requires Docker and an OpenRouter API key.

terminal — ~/your-project
>/plugin marketplace add brainmuxhq/brainmux
>/plugin install llmproxy@brainmux
$bmux init
$bmux config add-key OPENROUTER_API_KEY # hidden prompt — key not echoed
$bmux up
$bmux test
chat OK · deep OK · coder OK

FAQ

Questions, answered.

What is brainmux?

brainmux is LLM tooling for Claude Code. Its first tool, llmproxy, lets you run Claude Code on cheap alternate LLM models and delegate grunt work to them — so your Opus subscription quota goes to architecture and review, not busywork.

What is llmproxy?

llmproxy is a Claude Code plugin (CLI: bmux) that routes Claude Code to cheap OpenRouter models via local LiteLLM proxies. One OpenRouter key reaches thousands of models across providers like DeepSeek, Qwen, GLM, GPT and Gemini.

Does it use my Anthropic or Opus quota?

No. The cheap brains run pay-as-you-go on OpenRouter, on a separate meter that never touches your Anthropic subscription quota. Opus stays the orchestrator; the cheap brains do the volume.

What models can I use?

Any model on OpenRouter — DeepSeek, Qwen, GLM, Kimi, GPT, Gemini and hundreds more — with a single key. `bmux models` browses the live catalog by price, context and use-case, so you never guess a stale model slug.

How much does it cost?

llmproxy is free and MIT-licensed. You only pay OpenRouter's pay-as-you-go per-token price for the models you actually use — usually cents. Your Anthropic subscription quota is untouched.

How do I install it?

In Claude Code, run `/plugin marketplace add brainmuxhq/brainmux` then `/plugin install llmproxy@brainmux`. Then `bmux init`, add your OpenRouter key (`bmux config add-key OPENROUTER_API_KEY`), `bmux up`, and `bmux test`.

Do I need Docker?

Yes. Each brain runs as a local LiteLLM Docker container isolated by port, with one shared Postgres. brainmux generates the whole Docker Compose stack from a single brains.yaml file.