Last week we argued that buying more compute does not fix a memory problem. The missing piece is a router model, and this week Cloudflare shipped it.
They released Clef and Clef-flash, open source decision models built to return a bounded, structured answer, cheaply and fast. Their own description: “a model that produces bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required.”
We have been running that pattern for years. A small router model sits in front. It classifies the request, picks the tool, and decides whether the job is done or whether it has to escalate. The big model only wakes up when the work actually needs it.
What a router model changes
A router model is not a cheap substitute for your reasoning model. It is a gate.
Most of what hits a production system is not reasoning. It is triage. Which tool. Which queue. Is this safe. Does this even need a model. A small model handles that in a fraction of the memory and a fraction of the time, then hands the hard problem up.
We run it that way ourselves. A compact model takes the first look and escalates to a stronger one only when the request earns it. Classification stays cheap. Reasoning stays reserved for the work that justifies the cost.
That is the whole design. You do not need the smartest model in the world answering every question. You need the best reasoning model you can afford, and a router mean enough to protect it. Our InnoAI services are built around that split.
What Cloudflare measured
Cloudflare tested Clef on classifying website domains for its threat intelligence team. Clef took 2.2 seconds to fetch, render, and classify a site. Their fastest general model in the same workflow, gpt-oss-120b, took 4.7 seconds and returned only two classifications.
Those are Cloudflare’s numbers, published on the Cloudflare blog, so run them against your own traffic before you trust them. The direction is the part that matters. The specialist finished in under half the time and gave back more answers.
Read that against last week’s piece on why AI inference is a memory problem. A smaller model does not only answer faster. It moves fewer bytes, so it spends less of the bandwidth your decode path is already short on.
Put the router model to work this quarter
- Count the traffic. Split your requests into decisions and real reasoning. Most teams are shocked by the ratio.
- Put a small model in front. Let it classify, route, and refuse. Escalate only what fails that test.
- Measure cost and latency per request, not per token. A cheap token on the wrong model is still an expensive request.
- Trust your own numbers. A vendor benchmark is a starting point, not a plan.
Cloudflare did not invent this. They productized it, open sourced it, and put a number on it. If your stack still sends every request to the biggest model you own, you are paying reasoning prices for sorting mail.
Bring us a workload. We will show you which requests never needed the big model at all. Start with InnoCloud, or talk to us directly.



