Claude Code locks in a model the moment you press enter, so a bulk rename and a hard architecture question cost the same.

I built a hook that classifies each message in about 250 milliseconds and, for the cheap buckets, tells the agent to delegate to a sub-agent on a smaller model while hard problems stay with the strong model.

It is off by default, controlled by aggressiveness presets, costs about three cents per thousand messages, and fails open: if anything breaks, the message goes through untouched. Open source, with a README that states plainly what it saves and when not to turn it on.

README: two example messages, one routed to a cheap sub-agent, one kept on the main model, cost and latency shown.
README: two example messages, one routed to a cheap sub-agent, one kept on the main model, cost and latency shown.

View the code on GitHub

Back to Work