AIAI News Online

Aleph Alpha's Kolibri explained: a 78B model that computes like a 3.5B one

Aleph Alpha has released Kolibri, a German-English open-weight AI model trained in Europe. How its 384-expert design works, and where its own tests show gaps.

14 min at full depth9 sources

In 60 seconds

  • Aleph Alpha released Kolibri on 3 October: a German-English language model with open weights under Apache 2.0, trained on infrastructure in Germany and Finland and documented in a 189-page technical report.
  • Each of its 50 layers holds 384 small 'expert' networks and a router sends every token through only 6 of them, so 78.1 billion stored parameters cost about as much computation per token as 3.46 billion.
  • Regulated organisations get a documented model they can run on their own hardware, but every benchmark so far is Aleph Alpha's own, and its own tables show Alibaba's dense Qwen3.8 27B scoring higher overall in both languages.

In September 2024 the founder of the German AI company Aleph Alpha said that building a European language model was not, on its own, a business. On 3 October 2026 the company released one anyway: Kolibri, a model anyone can download, with a 189-page account of how it was made. It stores 78 billion parameters but uses only about 3.5 billion at a time, and that design explains both what Kolibri is good for and why its makers' own tables show a rival beating it.

For: Everyone

The plain-English version

An AI language model is, physically, a very large file of numbers called parameters. Most AI companies keep that file on their own servers and rent out access. Aleph Alpha has instead published the file (this is what "open weights" means) under a licence that lets anyone run it, including commercially. Kolibri, German for hummingbird, works in German and English. The company says it built the model in Germany and trained it on computers in Germany and Finland, "under European and German law". The intended customers are public authorities, manufacturers and aerospace firms that are not allowed, or not willing, to send sensitive documents to someone else's servers.

The clever part is how it keeps running costs down. Picture a 50-storey office building. On every floor sit 384 specialists and one generalist. Each piece of work travels up through all 50 floors, but on each floor a receptionist hands it to only six specialists, chosen for that particular piece, plus the generalist. You must rent the whole building and keep every specialist on the payroll. But any single job takes the time of just seven people per floor.

That is Kolibri. It has 78.1 billion parameters in total, and each fragment of text it processes passes through only 3.46 billion of them, about 4.4%. The rent is computer memory, which has to hold everything. The staff time is computation, which is what makes a model slow and expensive when many people use it at once. Kolibri pays more rent to save on staff time.

Aleph Alpha also trained it to admit ignorance. When the documents it is given do not contain the answer, it is supposed to say so instead of inventing something.

Two cautions. First, every test score so far comes from Aleph Alpha itself, and its own technical report shows a rival from Alibaba's Qwen family, with about a third as many parameters in total, scoring higher overall in both languages. Aleph Alpha's argument is that Kolibri needs far less computation per answer, not that it is the best. Second, "runs on your own hardware" means data-centre-class graphics processors with room for a 78-gigabyte file, not a laptop.

For: Curious

How it actually works

The problem. The report opens with the customer, not the technology: organisations that "process sensitive data under regulation" need a model that runs "on infrastructure that the organisation controls" and handles German as well as English. It adds a constraint that is easy to overlook: "the serving cost of the model limits the scale at which an organisation can deploy it."

The old way. In a conventional "dense" model, every parameter takes part in processing every token (a token is a chunk of text, roughly a word or part of one). Make the model larger and more knowledgeable, and every answer gets slower and costlier. The other option is to download an open model built elsewhere, and the strongest of those now come from Chinese labs, as Trending Topics notes.

The new idea, step by step.

  1. Many small experts, few used at once. A language model is a stack of layers. Each layer has two parts: attention, which lets a token gather information from other tokens, and a feed-forward block, which transforms each token on its own. Kolibri replaces that single block with 384 small ones, called experts, plus one shared expert that every token uses. A small "router" scores the experts for each token and picks the top six. This is a mixture-of-experts (MoE) design. It is not new: DeepSeek-V3 picks 8 of 256. Kolibri pushes the ratio further.

  2. Pay in memory, save on computation. Because only seven experts run per layer, each token costs about as much arithmetic as it would in a 3.46-billion-parameter model. But, as the model card puts it, "the full model must be held in memory even though only part of it is active at any time". The weights, stored at 8-bit precision, take about 78 GB. The minimum hardware listed is a single H200 or B200 GPU, or two 80 GB A100s.

  3. Size set by serving cost. Aleph Alpha trained three test versions with the same active size and different totals:

Total parameters (3.4B active in each) Quality in Aleph Alpha's tests Concurrent 256k-token requests on two H100 GPUs (estimated)
42.1B lowest of the three 33
82.3B better on all reported metrics 18
122.6B "only slightly further" improvement 3

The extra weights of the largest version crowd out the memory needed to serve several long requests at once. The blog says this "drove the decision", and that the chosen size also decodes 28% faster than the largest.

  1. Cheap long documents. In ordinary attention every token looks at every earlier token, so the work grows with the square of the document's length. In 40 of Kolibri's 50 layers a token looks only at the previous 512 tokens; every fifth layer looks at everything. Reading is mostly local, with a regular global check. The model accepts inputs of up to about a million tokens, though it was trained on sequences of up to 262,144.

  2. A tokeniser cut for German. German builds long compound words, which tokenisers designed around English chop into many pieces. Kolibri's splits along word parts: the blog's example is "Bundessozialgerichtes" becoming "Bundes", "sozial", "gericht", "es". The report says it needs 11.2% fewer tokens than GPT-5's tokeniser on German web text. Fewer tokens means less computation per page.

  3. Rewarding "I don't know". In training, the same question is posed over a document in three versions: intact, with unimportant sentences masked by a helper called Merlin, and with the crucial sentences masked by an adversary called Morgana. The model, Arthur, is rewarded for answering the first two and for abstaining on the third. In the blog's words, "Arthur does not know which one he's facing", so the only winning strategy is to check whether the text really supports an answer. The method comes from a paper that the model card lists as related work.

Why it works. What a model can store scales with its total parameters; what an answer costs scales mostly with the active ones. An MoE buys the first without most of the second, as long as the memory fits.

For: Practitioner

The deep dive

One block, fifty times

Kolibri is a decoder-only Transformer with 50 blocks of one shared design, a residual width of 2,560 and a 128,000-token vocabulary, per the technical report. Attention is grouped-query attention with 48 query heads, 4 key-value heads and head dimension 128. Every block has an MoE layer of 384 routed experts, each a SiLU-gated feed-forward network with hidden dimension 512, plus one shared expert. The router computes logits ZZ, adds a per-expert bias β\beta that affects selection only, and weights the chosen experts by sigmoid scores that are not renormalised:

It=TopK(Zt,:+β, 6),yt=fshared(xt)+∑e∈Itσ(Zt,e) fe(xt)\mathcal{I}_t=\text{TopK}(Z_{t,:}+\beta,\,6),\qquad y_t=f_{\text{shared}}(x_t)+\sum_{e\in\mathcal{I}_t}\sigma(Z_{t,e})\,f_e(x_t)

The experts are tiny. Assuming three weight matrices per expert (the optimiser table lists up, gate and down projections), each holds about 3.9 million parameters, and the 19,200 routed experts together hold roughly 75.5 billion, about 97% of the model (our arithmetic). Almost everything Kolibri stores is expert weight that a given token never touches.

The closest prior designs are less sparse. DeepSeek-V3 activates 8 of 256 routed experts of hidden dimension 2,048 and normalises the selected scores; 37 billion of its 671 billion parameters are active, about 5.5%. Kolibri's unreleased predecessor, Kolibri Origin, routed to 8 of 128 experts, with 10.7% active. An ablation at about 80 billion total parameters, after 327.6 billion training tokens, compared three shapes:

Configuration Width Experts Top-k Expert hidden Loss Accuracy
Kolibri 2,560 384 6 512 1.847 38.0%
Finer (Qwen3-Next's expert count) 2,048 512 12 512 1.848 37.6%
Coarser 2,048 256 6 1,024 1.860 35.6%

Reported by Aleph Alpha, technical report Table 32.

Keeping 384 experts busy

Sparse routing fails if a few experts absorb most tokens. The report tracks the maximum relative overload:

MaxVio(B)=max⁡e[ce(B)K∣B∣/E−1]\text{MaxVio}(\mathcal{B})=\max_e\left[\frac{c_e(\mathcal{B})}{K|\mathcal{B}|/E}-1\right]

where cec_e counts the tokens sent to expert ee. Kolibri uses two mechanisms the report credits to Neitemeier et al. (2026). Exact Quantile Balancing sets each expert's bias to the quantile of its adjusted logits that would give it exactly its fair share, TK/ETK/E tokens. The difficulty is that the batch is spread across hundreds of GPUs. Earlier approaches averaged per-GPU quantiles or approximated with a 1,000-bin histogram; EQB recovers the exact value for 16-bit logits with a two-pass radix selection (a 256-bin histogram over the high byte, then over the low byte), costing two all-reduces of 256E256E integers per layer per step, independent of batch size. Load-Error Injection balances each local microbatch by adding every expert's relative load error to the gradient of its routing score, with strength λ=10−5\lambda=10^{-5}. Global MaxVio averaged 0.31 across layers during pre-training.

The report then documents where this fell short. In the first two layers the bias overrides the router almost completely: the experts actually used differ from the router's own top six 98.5% of the time, which is what a random choice would give (1−K/E1-K/E). Swapping those layers' routed experts for random ones, or silencing them, changes the loss by an amount "indistinguishable from zero". A fix attempted during fine-tuning brought "no clear downstream improvement", and the report states: "Resolving this issue remains future work." A shorter ablation had pointed the other way; the behaviour "only emerged after longer training durations".

Attention that mostly looks nearby

Forty blocks use sliding-window attention over the 512 preceding tokens with rotary position embeddings. Ten use full attention with no positional encoding at all, which the team found improved long-context results. For serving, the point is the key-value cache: in windowed layers it stops growing, and the report says most long-context memory traffic comes from the ten full-attention layers. The window size was a judgement call. At the larger proxy scale a 128-token window scored 48.0% on pre-training benchmarks against 47.8% for 512, but at 256,000 tokens of context the 512 window was 6.7 points better. The model card describes the 1,048,576-token maximum as reached by extrapolation and recommends contexts of 262,144 tokens or fewer.

Training, tokeniser and cost

Stage Tokens Sequence length Duration GPU-hours
Pre-training 20T 16k 21 days 392k
Mid-training 3.44T 64k 5 days 90k
Long-context extension 201B 256k 13 hours 10k

Reported by Aleph Alpha on the model card; 768 Nvidia B200 GPUs.

The model card puts compute at 6.4×10236.4\times10^{23} FLOPs and energy at an estimated 950 MWh for these three stages, excluding fine-tuning and reinforcement learning. Attention and expert matrices were optimised with Muon, the router and output head with AdamW, embeddings and norm scales with Adam. The learning rate stayed constant after a 100-billion-token warm-up, and the base model is the average of the last 20 checkpoints, 10 billion tokens apart, in place of a cooldown.

The tokeniser method, UniBPE, builds merges bottom-up like byte-pair encoding but ranks each candidate by how much it lowers the Unigram loss L=−∑tctlog⁡(ct/N)\mathcal{L}=-\sum_t c_t\log(c_t/N), not by raw pair frequency; the last 7,900 merges use standard byte-pair encoding. The report says it reaches 4.90 bytes per token on German web text against 4.35 for GPT-5's tokeniser, and 4.58 against 4.67 on English.

Post-training is supervised fine-tuning (537 billion tokens, per the report's summary) followed by 1,000 reinforcement-learning steps drawing on more than 1.2 million tasks. Rewards are verifiable and lie in [0,1][0,1], except in the Merlin-Arthur environments, where they are signed in [−1,1][-1,1]; a sentence's importance is estimated from how far the model's probability of the correct answer falls when that sentence is masked. With German-language environments in the mix, the report says, accuracy on German AIME 2026 rose from 75% to 91% over the run.

The numbers, and who reported them

Every figure below was measured by Aleph Alpha in its own evaluation setup (technical report Tables 28 and 29), with Kolibri at high reasoning effort. None has been independently reproduced.

Benchmark Kolibri (78B, 3.46B active) Qwen3.5 35B-A3B Qwen3.6 35B-A3B Nemotron 3 Super 120B-A12B Qwen3.8 27B (dense)
Overall, English 75.5 74.7 71.4 73.0 80.2
Overall, German 70.8 69.8 67.3 67.9 79.9
AIME 2025 (English) 96.9 88.1 84.6 91.7 97.9
GPQA Diamond (English) 84.3 83.8 83.4 78.0 89.2
LiveCodeBench v6 85.9 77.8 82.5 82.0 93.8
SWE-Bench Verified 66.4 71.6 73.8 60.2 72.6
AA-Omniscience accuracy 14.8 22.0 19.5 26.7 17.5
AA-Omniscience non-hallucination rate 44.0 11.1 56.7 13.9 67.3
RGB closed-book 51.0 81.0 79.0 93.0 73.0

On throughput, Aleph Alpha estimates that Kolibri decodes 2.7 times as much text per GPU as Kolibri Origin on eight B200 GPUs, alongside a gain of more than 21 points on the overall score.

For: Everyone

Why it matters

Everyday users. Few people will chat with Kolibri directly. If it is adopted, it will sit behind the document systems of the public bodies and regulated firms it was built for. What reaches them is a model that reasons in German (on 99% of German AIME 2026 prompts after training, the report says) and is trained to decline when the supplied documents do not support an answer.

Developers and builders. The weights are Apache 2.0, served through vLLM with an Aleph Alpha plugin, with a per-request reasoning-effort setting. The model page already lists several quantised versions. Early hands-on reports are anecdotal: one Hacker News commenter reported "around 170 tkn/s on fp8" on an RTX Pro 6000 setup. Engineer Tejas Kumar, who notes a friend works at Aleph Alpha, ran the tokeniser over the German constitution and counted 35,190 tokens against 41,482 for GPT-5's. Kumar's verdict is that Kolibri is "the pick when German text and your own hardware both matter" and the wrong pick "for a coding agent".

Companies. For a regulated buyer the product is provenance as much as performance. The model card documents filtering, redaction of personal data and decontamination, and says Aleph Alpha is a signatory of the EU's GPAI Code of Practice. The blog's argument is that a model that knows to abstain "is the difference between a pilot and a deployment". The report also gives planning figures: an estimated 18 concurrent 256k-token requests on two H100s.

The field. The report prints ablations and failures alongside results, which is useful whatever one thinks of the checkpoint. The release also marks a reversal. In September 2024 Jonas Andrulis, then Aleph Alpha's chief executive, told Bloomberg, as TechCrunch reported: "Just having an European LLM is not sufficient as a business model."

Two second-order effects stand out. First, sparsity is a bet on hardware. The memory penalty that ruled out the 123-billion variant on H100s nearly vanishes on newer GPUs: on two B300s the report estimates 142 concurrent long requests for it against 157 for the roughly 80-billion design. More memory per GPU pushes the best design towards larger, sparser models. Second, the report argues that one of its most durable results is "neither a checkpoint nor a training pipeline" but "a team in Europe that has trained two Kolibri generations in a short amount of time". That frames sovereignty as a capability to build, not a file to own.

For: Critical

What to be skeptical of

All the scores are the vendor's. Aleph Alpha ran every model in its own evaluation setup; the report itself calls its base-model results a comparison "under one protocol, not a reproduction of published numbers". The five "industry" benchmarks are in-house proxies, and the report says two of them, with 11 and 26 items, "are too small to show a trend".

The comparison set is dated. Trending Topics points out that stronger open-weight models, including Z.ai's GLM-5.3 and Moonshot's Kimi K3, were not tested. Its estimate that Kolibri would score 15 to 20 on the Artificial Analysis index, behind at least 20 open-weight models, is an extrapolation, not a measurement. On Hacker News, one commenter called the absence of Qwen3.8 Flash "pretty striking".

A model a third of Kolibri's size wins in Aleph Alpha's own table. Qwen3.8 27B has 27 billion parameters in total, though all of them are active for every token. As another commenter put it, "Qwen3.8 27B beats Kolibri 79.9 vs 70.8 in German in Kolibri's harness on Kolibri's benchmark." A commenter who said they worked on Kolibri at Aleph Alpha replied that it is a trade-off: "Kolibri needs less compute per token but more memory, while Qwen3.8 27B needs far less memory and more compute per token."

The hallucination claim is relative. On AA-Omniscience, when Kolibri does not get a question right it abstains or gives a partial answer, instead of guessing wrong, 44.0% of the time, against 15.0% for its predecessor. But Qwen3.6 35B-A3B scores 56.7% on the same measure and Qwen3.8 27B 67.3%. On the RGB closed-book test Kolibri is last of the twelve models compared. One Hacker News user reported that it invented an album and year for a made-up song.

Long context is uneven. As a base model Kolibri scores 67.9 on RULER at 128k tokens against 89.9 for Qwen3.5's base model; it has the best score only at one million tokens, beyond its own trained window.

"Sovereign" has asterisks. The training ran on Nvidia GPUs. The model card says the training data "also contains material generated with Chinese language models", and the report says a Z.ai model was used to synthesise retrieval tasks. The licence covers weights and configuration, not training code. And Aleph Alpha has signed a binding agreement to merge with Canada's Cohere, with the combined company to operate under the Cohere name, Trending Topics reports.

For: Everyone

What to watch next

  • Independent benchmarks. Trending Topics had to estimate because Kolibri had no Artificial Analysis score at launch. A listing would test its 15-to-20 figure.
  • A head-to-head with Qwen3.8 Flash, the comparison commenters asked for, run by someone other than either vendor.
  • Quantised versions. The release is already 8-bit. Whether quality holds at lower precision decides if Kolibri can run on less than 78 GB.
  • The next Kolibri. Pre-training of Kolibri Origin finished on 11 June and of Kolibri on 11 September, per the blog. The report names better German reasoning, fewer hallucinations and more European languages as next steps, and the blog says the team is "contemplating scaling up".
  • The first-two-layers problem. A fix, or more detail from the Exact Quantile Balancing work the report cites, would show whether very sparse routing is sound at this scale.
  • The Cohere merger. Whether and when it closes, and what the combined company does with the Kolibri line.

For the opposite release strategy, a frontier model held behind an access gate, see our explainer on Gemini 4 Argon.

Check your understanding

Pick an answer — you'll see why right away.

1. Kolibri has 78.1 billion parameters, of which 3.46 billion are 'active' per token. What does that mean for someone running it?

2. Aleph Alpha's experiments found that a 123-billion-parameter variant scored slightly higher than the roughly 80-billion design. Why did it ship the smaller one?

3. In Aleph Alpha's own results table, the dense Qwen3.8 27B scores higher overall than Kolibri in English and German. How can Aleph Alpha still say Kolibri is on the 'Pareto frontier'?

4. How does the Merlin-Arthur training teach Kolibri to say it does not know?

Glossary

Open-weight model
An AI model whose trained parameters are published so that anyone can download and run it on their own hardware.
Parameter
One of the billions of numbers inside a model that are adjusted during training and together determine its behaviour.
Mixture of experts (MoE)
A design in which each layer contains many small sub-networks and a router sends each token through only a few of them.
Active parameters
The subset of a mixture-of-experts model's parameters that actually take part in processing a given token.
Token
A chunk of text, often a word or part of a word, that a language model reads and writes one at a time.
Tokeniser
The component that splits text into tokens before the model sees it.
Context window
The maximum amount of text, measured in tokens, that a model can take into account at once.
Sliding-window attention
A cheaper form of attention in which each token looks only at a fixed number of preceding tokens instead of the whole input.
Pareto frontier
The set of options for which no alternative is better on one measure without being worse on another, here quality and serving cost.
Abstention
A model declining to answer, for example by saying the supplied documents do not contain the information.

Questions people ask

What is Aleph Alpha Kolibri?

Kolibri is a German-English language model released by the German company Aleph Alpha on 3 October 2026. Its weights are published on Hugging Face under the Apache 2.0 licence, and it is a mixture-of-experts model with 78.1 billion parameters, of which 3.46 billion are active per token.

What hardware do I need to run Kolibri?

The model card lists a minimum of two 80 GB A100 GPUs, two H100s, or a single H200, B200 or B300, because the weights take about 78 GB of GPU memory. It is served with vLLM and an Aleph Alpha plugin. It will not run on a typical laptop, although the model page already lists quantised versions.

Is Kolibri better than Qwen or Mistral models?

By Aleph Alpha's own tests it scores above the mixture-of-experts models it was compared with, including Qwen3.5 and Qwen3.6 35B-A3B and Mistral Small 4, on the overall averages. The same tables show the dense Qwen3.8 27B ahead of it in both English and German, and newer open-weight models were not tested. As of 4 October 2026, no independent benchmark had been published.

Is Kolibri open source?

The weights and configuration files are released under Apache 2.0, which permits commercial use. The model card states that the licence does not extend to the underlying code, architecture or training methods, so it is more accurately called open-weight than open source.

Does Kolibri work in languages other than German and English?

It was built as an English-German model, with German making up more than a fifth of its training data. Aleph Alpha's report names extending to other European languages as an area of interest for future work, so other languages are not a design target of this release.

What does 'sovereign' mean for Kolibri, given the Cohere merger?

Aleph Alpha uses the word for two things: that it built and trained the model in Europe under European and German law, and that customers can run it on infrastructure they control. Aleph Alpha has signed a binding agreement to merge with Canada's Cohere, as Trending Topics reports, which has led commenters to question the label; the published weights remain usable under Apache 2.0 either way.

Discussion

  1. Loading comments…

Sources

  1. Kolibri Has Landed: A Sovereign Open-Weight Model — Aleph Alpha · official announcement
  2. Kolibri: A Sovereign European Model on the Pareto Frontier — Aleph Alpha · paper
  3. Aleph-Alpha/Kolibri-1 — Hugging Face · docs
  4. Aleph Alpha's Sovereign A.I. Model Kolibri Is No Match for the Open-Weight Leaders — Trending Topics · analysis
  5. Aleph Alpha Kolibri: How the Sovereign German LLM Works — Tejas Kumar · analysis
  6. Kolibri: A Sovereign Open-Weight Model (discussion) — Hacker News · analysis
  7. Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models — arXiv · paper
  8. DeepSeek-V3 Technical Report — arXiv · paper
  9. German LLM maker Aleph Alpha pivots to AI support — TechCrunch · news

How this was made: researched and written by an AI model (Claude) from the primary sources listed above, then checked claim-by-claim against those sources in a separate AI fact-check pass. Spotted an error? Email [email protected] and we correct it publicly. Our process.