GotMoo?
The router for Claude Code. Local-first. Learns forever.
Spawns agents safely by default.
Your GPU, your subscriptions, your local models — you're already paying for a powerful AI stack. But Claude Code defaults to Opus for everything, even renaming a variable. Mooter maps your full environment and routes every prompt to the optimal model — 101/123 classified prompts went to a local or cheap tier. No tokens are logged, so there is no measured dollar figure — and we publish none. See the benchmark *
For a vibe coder on a Max plan: renames, commits & explains run local (free); debugging & refactors go cloud — so the free/cloud split shifts with the work the day brings, not with a fixed ratio. Estimate yours →
Honest numbers
On 2026-08-23 an audit traced every savings figure this project had published. Five were in circulation and they contradicted each other; none survived. The cause is structural and worth stating plainly: no telemetry file in this project records token counts. Without tokens there is no measured cost, and without a measured cost there is no measured saving — in any unit.
What is measured. The router recommended a local or cheap tier for 101 of 123 classified prompts (82.1%). In those same sessions, of the 3225 recorded executions, 3193 ran on Opus and 1 ran locally.
That gap — what the router recommended against what actually ran — is the honest state of the project, and closing it is the current work. No tokens are logged, so there is no measured dollar figure — and we publish none. A percentage published on top of this would describe a product that does not exist yet.
Same prompts. Comparable quality on routine tasks.
Two bills, line by line.
The exact same six prompts, streamed into both at once. The hard one — the schema migration — stays on Opus on both sides. Mooter only routes down when quality holds.
Your GPU, your subscriptions, your local models — you're already paying for a powerful AI stack. But Claude Code defaults to Opus for everything, even renaming a variable. Mooter maps your full environment and routes every prompt to the optimal model: comparable quality on routine tasks, a fraction of the spend.
For a vibe coder on a Max plan: renames, commits & explains run local (free); debugging & refactors go cloud — so the free/cloud split shifts with the work the day brings, not with a fixed ratio. Estimate yours →
Why a local model is good enough
See the benchmarks →Quantization
Q4_K_M shrinks a 120 GB model to ~18 GB. Quality stays (~98%), speed goes up, and it runs free on the GPU you already own.
Learn more →LoRA / DoRA
Adapter layers can fine-tune a base model for one codebase without re-training it. Not shipped yet — the explainer covers how it works and where Mooter is headed.
Learn more →Hardware match
mooter probes your GPU/CPU and pulls only the models that fit your VRAM. No config, no guesswork — it just fits your machine.
Learn more →Never wake up lost
How it stays honest →The real cost of agentic coding is not tokens — it is opening ten terminals and having no idea where you left off. Mooter closes every session with an honest handoff: what landed, what did not, and the one thing to do next. So you recover your time instead of drowning in context.
Honest summaries
After every session Mooter writes what actually happened — and what is still unlanded. It never says “clean” when there is work waiting to ship. No false green.
Auditable memory
Every decision gets recorded — which tier, why, what it cost. The local GPU does the writing, so the record itself costs $0. Nothing is taken on trust.
The time back
Glance and know what to do next. No re-reading diffs, no archaeology across terminals. You pick up exactly where you left off — that is the time you get back.
*Illustrative — real community numbers appear once devices start phoning home.