skip to content
The Weighted Average

Models & Open Source

Kolibri's FP8 Weights Fill About 98% of One H100

Kolibri's 78 GB FP8 weights nearly fill an 80 GB H100, and Aleph Alpha requires two. Size the serving system before buying its sparse-model promise.

Server racks with orange and blue network cables
Server racks with orange and blue network cables. Photograph by Kevin Ache

Size the deployment before booking the GPU: Aleph Alpha released Kolibri on October 3 with approximately 78 GB of FP8 weights. Against NVIDIA’s 80 GB H100 SXM specification, those weights occupy 97.5% of nominal memory before serving overhead; Aleph Alpha’s minimum is 2 H100 GPUs, not one.

Sparse computation still needs somewhere to live

The arithmetic is deliberately narrow. Divide the model card’s approximate 78 GB footprint by NVIDIA’s 80 GB capacity: 78 ÷ 80 = 97.5%. Subtraction leaves approximately 2 GB, before runtime allocations and the context cache. This is a comparison of published specifications, not a measurement of free memory, a successful loading experiment, or a claim that a single H100 can serve Kolibri. The published two-card minimum is the operational instruction.

That distinction changes the first procurement question. A team evaluating the release should request the supported configuration for its intended workload, rather than ask whether the weight file is smaller than a card’s label. Apparent fit is only the beginning of a capacity review. The resource that needs approval is a working endpoint under load, with room for the requests it must keep in flight.

The FP8 model card lists 78 billion total parameters but only 3.46 billion active per token. The active count describes sparse computation; it does not shrink the resident model to that count. A procurement specification needs the checkpoint, hardware and serving configuration together. The active count cannot replace that specification.

Precision changes the purchase too. The separate BF16 checkpoint has approximately 156 GB of weights and lists four H100s as its minimum. That is a different artifact, with different deployment guidance. A benchmark, infrastructure quote, or internal approval that says only “Kolibri” leaves an avoidable ambiguity. Record the checkpoint and precision beside the hardware before comparing results or authorizing a reservation.

There is a practical route to testing. The official inference repository supplies a vLLM plugin and currently supports vLLM 0.29. Its reference FP8 command uses an FP8 context cache, while its BF16 instructions change the checkpoint and omit that cache option. This is an integration dependency to pin and test, not an invitation to assume an existing inference image already behaves identically. Preserve a working environment before changing it.

Kolibri’s appeal extends beyond its memory footprint. The launch account describes a German–English model released under Apache 2.0, with training infrastructure in Germany and Finland. That makes it a candidate for organizations seeking a locally operated language model. It does not settle an application’s data-handling obligations, quality threshold, or operating bill. Control is valuable when the team has a concrete reason to exercise it and the capacity to maintain the resulting service.

Test the request that will actually arrive

Context is the next constraint to write down. The model card advertises 1,048,576 tokens, but recommends 262,144 or fewer for efficient serving and complex tasks. Those are different promises. Start the pilot with the documents the application actually needs, then expand only when the measured result warrants it. A maximum accepted length is not a guarantee that every long request will meet the required accuracy or latency.

The strongest case for adoption is therefore specific: a German- or English-language document workflow where deployment control matters and the team can qualify its own endpoint. The weakest case is replacing an incumbent solely because a sparse model sounds inexpensive. The retrieved release materials do not establish the buyer’s hourly hardware tariff, utilization, support expenditure, or cost per accepted answer. No honest break-even traffic figure follows from memory capacity alone.

Ask the pilot to answer a purchasing question. Run representative documents through the intended configuration and record accepted outputs, request duration, peak memory, and review effort. Include a busy interval rather than testing only isolated requests. If the service needs spare capacity to meet its response target, charge that capacity to the comparison. If human correction remains substantial, count that work too. These are proposed acceptance measurements, not results this article claims to have obtained.

Quality deserves an equally explicit test. Aleph Alpha’s launch post reports evaluations using its own harnesses and discusses grounding and abstention training. Treat those as vendor evidence for choosing test cases. Include questions whose answers are present in the documents and questions whose answers are absent; reviewers should distinguish a justified refusal from an unsupported answer. An application that rewards every fluent response can miss the behavior that makes a document assistant useful.

The archive’s Arcee analysis separated active parameters from the full deployment footprint. Kolibri turns that distinction into a concrete hardware check. The lesson is not to reject sparse models: computation savings can be real while deployment requirements remain substantial. The correct response is to price and measure the configuration the vendor actually supports, then compare its accepted work with the current service.

Today’s ChatGPT advertising lead separates a reported outcome from the evidence needed to attribute it. Apply the same discipline here. Lower resource use, acceptable answers, and deployment control are separate claims that need separate evidence. A fast short request does not prove economical long-document service; a permissive license does not demonstrate acceptable answers; a large context ceiling does not establish usable throughput.

Pilot Kolibri where language fit and operational control justify the work, and budget against a supported configuration. Broader adoption becomes rational if the tested endpoint meets quality and latency requirements at a lower complete operating cost, or if its control benefits justify a measured premium. Delay the switch if serving overhead, review burden, or integration work erases the advantage. The purchasing test is a working service, with its overhead measured and its useful output accepted.

Sources