Aleph Alpha released Kolibri today, and the number that matters is not the big one. Kolibri is an English-German Mixture-of-Experts model with 78 billion parameters in total and 3 billion active. A mixture of experts model keeps the whole network on the machine and only wakes the slice the token actually needs.
Their line, from the announcement: the model’s “small and efficient size provides our customers with flexibility to run it efficiently on-premise, without sending internal data to third-party inference services.”
That is the argument we have been making. You do not serve the whole model. You serve the part that earns it.

Why a mixture of experts model changes the bill
A dense model reads every parameter for every token. That is the memory wall. Weights move out of memory whether the token needed them or not, and the GPU sits waiting on bandwidth you already paid for.
A mixture of experts model refuses that. Kolibri holds 78 billion parameters and activates 3 billion of them. Aleph Alpha says it sits on the Pareto frontier for quality against serving cost, in both English and German: none of the models they compared delivered more quality at the same cost, or the same quality for less. Those are their benchmarks. Read the tech report before you trust the chart.
The direction is the part you can use. Active parameters are what you pay to move. Total parameters are what you keep on disk.

This is the same cut as routing
We wrote this week that a small router model should decide what your big model sees. Kolibri makes that cut inside a single set of weights.
A router sends the easy request to a small model and escalates the hard one. A mixture of experts model sends each token to a small set of experts and leaves the rest cold. Different layer, same discipline. Do not light up capacity the request did not earn.
Stack the two and the waste gets hard to justify. A router in front keeps classification, tool choice, and safety checks off the reasoning model. The reasoning model, if it is built this way, keeps most of its own parameters idle. You pay for the thinking, not for the parade.
On-premise is the point, not the footnote
Kolibri is Apache 2.0. The weights are on Hugging Face, context runs to 1 million tokens, and Aleph Alpha built it for regulated work: public administration, industrials, aerospace. They trained it to abstain, to say it does not know when the answer is not in the documents.

Sovereignty is their word for it. The practical version is simpler. If the data cannot leave the building, the model has to live in the building, and a 78 billion parameter dense model is a brutal thing to host. Three billion active is a different conversation. It fits hardware a company can actually own.
We have run production inference long enough to watch the other path fail. Teams rent a frontier model, send every request to it, and then discover the invoice and the data-handling review arrive in the same quarter. A model you can run yourself, that only spends a fraction of its weights per token, is how you get out of that.
What to do with this
- Ask what is active, not what is big. A parameter count with no active count is a marketing number.
- Put the router back in front. Mixture of experts saves you inside one model. It does not stop you sending trivia to it.
- Decide what must stay on-premise. If the answer is the documents, the model has to be one you can host.
- Read the tech report. The Pareto chart is Aleph Alpha’s. Your traffic is the only benchmark that counts.
Kolibri will not be the last model shaped this way. It is the one that made the case in public, with the weights attached. The teams that keep buying the biggest dense model they can rent are paying for parameters that never touch the token.
If your inference still moves every weight for every request, bring us the workload. We will show you what a mixture of experts model changes, and what a router in front of it changes again.


