Picture a massive library with 78,000 books, but when you ask a question, a smart librarian only walks over to the 3,000 books that actually have your answer, not the entire library.
That's essentially what Aleph Alpha just built with Kolibri. The model totals 78.1 billion parameters, but its Mixture-of-Experts (MoE) architecture only activates 3.46 billion parameters per token.
The result is the power of a big model at a compute cost closer to a much smaller one. It's the difference between lighting up every room in the house versus just the one you're standing in.
Kolibri is specialized for English and German, and it comes with a massive 1-million-token context window, meaning it can hold onto very long conversations or documents without losing track of details.
There's also a neat feature called per-request reasoning effort, basically letting you tell the model how hard to think on a given task, like telling an employee 'spend more focus on this one than the last.'
The weights ship in FP8 under an Apache 2.0 license, and they run on a single B200 or H200 GPU, no server farm required.
This points to a bigger industry shift: the race isn't just about who builds the biggest model anymore. It's about who builds the most efficient one that still performs under real-world resource limits.
🔗 Original Source & Reference: https://www.marktechpost.com/2026/10/04/aleph-alpha-releases-kolibri-a-78-1b-open-weight-english-german-moe-model-with-only-3-46b-active-parameters/
Published automatically via FeedMind AI Content Pipeline.

Top comments (0)