Ground Truth.
AI, checked against the source.

News · 2026-10-04

Aleph Alpha releases Kolibri, a German-English model with a 78 GB weight footprint

Aleph Alpha released Kolibri on October 3, 2026, giving organizations a German-English model they can download and run in an environment they control. Its sparse expert architecture reduces computation per word without removing a substantial hardware burden: the published eight-bit weights occupy about 78 GB, and the developer’s minimum configurations use data-center graphics cards.

Key facts

The important purchase decision is control. An organization handling German contracts, public records, or internal documents can evaluate the released model without sending inference inputs to a remote model provider. That changes where operational authority sits. It does not decide how well the model handles the organization’s particular documents, nor whether its deployment meets every legal obligation.

Aleph Alpha’s Hugging Face model card makes the practical trade clear. Kolibri uses a mixture of experts: a routing mechanism selects a small subset of specialized components for each token rather than activating the entire model. Imagine a large professional firm that sends each case to six specialists. The consultation can be efficient, but the firm still needs offices for all its specialists. Likewise, sparse computation does not make the inactive expert weights disappear from memory.

The released low-precision weight footprint is approximately 78 GB across 32 weight-file shards. The higher-precision variant is about 156 GB; its separate running-memory requirement was not established in the dossier. Disk size is separate from the graphics-card memory needed to run inference. For the eight-bit release, Aleph Alpha lists minimum configurations that include two A100 cards with 80 GB each, two H100 SXM5 cards, or one H200, B200, or B300. These are published configurations, not a claim that a 78 GB memory allocation is sufficient for every workload. Working memory and the model’s conversation cache come on top of the weights. No official consumer-card or four-bit requirement was established in the dossier.

The long-context headline also needs its qualification. The card distinguishes native training length from an extrapolated maximum of 1,048,576 tokens. The 189-page technical report reports encouraging results on a particular long-context test beyond training length. That shows performance under that test’s conditions, not dependable retrieval from every million-token document collection. For complex tasks and serving efficiency, the company recommends remaining within the native context.

This is a substantive engineering disclosure. The report covers architecture experiments, staged training, evaluation, infrastructure, and data review. It describes a pipeline that enabled many controlled comparisons and acknowledges that an earlier training generation was restarted after a data-shuffling bug. Such details make the report more useful than an announcement alone: they expose decisions and failure modes that potential users can interrogate.

The company says training used 768 NVIDIA B200 graphics processors in Germany and Finland, with “no foreign control.” That is Aleph Alpha’s description of the training arrangement. It is not an independent audit of infrastructure ownership or an assurance about permanent corporate governance. Its Cohere combination agreement makes future governance a separate question, contingent on transaction terms and closing.

Data and licensing deserve the same precision. The card’s Apache 2.0 grant applies to the published weights and configuration files; it does not license everything used to create the model. The report describes checks for licenses, restrictions, and ignored opt-outs in third-party datasets. That process is meaningful disclosure, but it does not prove comprehensive source-level rights clearance. The corpus is not described as exclusively European-licensed material.

The strongest performance objection is inside Aleph Alpha’s own evidence. Its broader evaluation table gives Qwen3.8 27B higher overall English and German averages than Kolibri. That undercuts a universal best-model interpretation of launch highlights, while leaving room for advantages on individual tasks or deployment criteria. Hacker News participants questioned both comparison choices and the memory burden; Tejas Kumar’s explainer helps explain the architecture but is not an independent broad benchmark.

Kolibri therefore adds a concrete option to the sovereign-government AI discussion. The useful next step is a controlled test on real German-English workloads, with the intended context lengths, memory configuration, and comparison models. Downloadability establishes deployment choice. Whether that choice is worth operating depends on measured quality, cost, and the organization’s actual control requirements.


Primary source, verified: read the paper →

Key questions

How much storage and GPU memory does Kolibri need?

The published eight-bit weights occupy approximately 78 GB on disk. Aleph Alpha lists minimum configurations including two A100 cards with 80 GB each or a single H200; serving also needs memory for working state.

Is Kolibri’s million-token context native?

No: its native context is 262,144 tokens, while 1,048,576 tokens is extrapolated. Aleph Alpha recommends staying at or below the native length for complex tasks and serving efficiency.

Does sovereign mean independently certified as legally compliant?

No: self-hosting gives customers deployment control, while data-rights and compliance statements remain Aleph Alpha’s own disclosures.
Cite this

APA

Ground Truth. (2026, October 4). Aleph Alpha releases Kolibri, a German-English model with a 78 GB weight footprint. Ground Truth. https://groundtruth.day/news/aleph-alpha-kolibri-sovereignty-meets-the-memory-budget.html

BibTeX

@misc{groundtruth:aleph-alpha-kolibri-sovereignty-meets-the-memory-budget,
  title  = {Aleph Alpha releases Kolibri, a German-English model with a 78 GB weight footprint},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/aleph-alpha-kolibri-sovereignty-meets-the-memory-budget.html}
}

Topics: open-weights · sovereign-ai · mixture-of-experts · german · inference

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.