skip to content
The Weighted Average

Models & Open Source

Reflection Beam: the Beta Migration Budget

Beam's beta has 25% of the context in a documented GLM-5.2 setup. Check endpoint behavior and reasoning budgets before switching agents.

An aisle between rows of metal equipment racks beneath exposed ceiling beams
An aisle between rows of metal equipment racks beneath exposed ceiling beams. Photograph by İsmail Enes Ayhan

Reflection’s October 5 Beam announcement opens an early-access waitlist while promising Apache 2.0 weights later this month, making a bounded API trial the immediate decision. The beta’s documented context budget is 25% of the GLM-5.2 configuration recorded by vLLM—a migration constraint worth checking before treating Reflection’s compute-efficiency claim as a production saving.

The model promise meets the request ceiling

Beam is a sparse model with 501 billion total parameters and 23 billion active per token, according to Reflection’s launch disclosure. The company positions it for coding, reasoning, and agentic work, and says its reasoning results use less inference compute than GLM-5.2. That comparison estimates generation computation; the accompanying methodology excludes prompt prefill, context-dependent attention, and serving overhead. It is not a measured customer bill.

The immediately usable contract is narrower than the eventual release promise. Reflection’s model documentation identifies a combined input-and-output ceiling of 262,144 tokens, with maximum output of 131,072 tokens. The page explicitly warns that the context window may change during the beta. Its example metadata is a published specification, not evidence that this publication has authenticated to the service or independently stress-tested that limit.

Compare that ceiling with the vLLM project’s GLM-5.2 recipe, which sets context_length to 1,048,576. The arithmetic is 262,144 ÷ 1,048,576 × 100 = 25%. A request assembled around that particular GLM configuration therefore faces a ceiling one quarter as large when moved to the documented Beam beta. This compares two recorded configurations. It does not establish the final Beam checkpoint’s intrinsic context capacity, a performance ranking, or the capacity of every hosted GLM service.

The practical implication is to qualify requests before qualifying savings. Repository snapshots, conversation history, instructions, tool descriptions, and tool results compete with generated output inside the same request budget. A team planning to move long-running agents must inspect actual request sizes and its truncation or summarization policy. A model can be attractive on short tasks while an existing orchestration strategy makes it unusable on the workload that matters most.

Even the output maximum is not a free addition to the window. Reserving the full 131,072-token output allowance leaves 131,072 tokens for input under the documented combined ceiling. That is a budget calculation, not a recommendation to reserve that much output on every request. Excessive reservation can displace useful evidence; insufficient reservation can interrupt the work. The right allocation depends on the task’s acceptance criteria and observed generation behavior.

This is the same distinction the archive drew when Atria’s hosted API imposed a concrete request budget despite broader model claims. Beam deserves evaluation on the interface available to the buyer today. Announced weights and a future serving stack cannot yet resolve how a particular production deployment will behave under load.

Compatibility carries its own invoice

Start the pilot with a protocol inventory. Reflection’s compatibility documentation supports Chat Completions and Models but excludes Responses and other endpoint families. A client limited to supported operations may need only a base-URL and key change. A workflow built around Responses needs additional integration work. The familiar SDK name is not evidence that the entire application contract transfers.

The less obvious difference is more consequential: the same documentation says stop is accepted but has no effect. If a downstream parser depends on generation stopping at a delimiter, an apparently successful request can violate its assumptions. Test the returned content and termination behavior, not just HTTP success. The service also accepts text input only, so a coding workflow that depends on direct screenshot inspection needs a separate visual path.

Then inspect the reasoning budget. Beam’s reasoning guide lists five effort levels and makes reasoning mandatory, with medium applied when effort is omitted. Reasoning tokens consume the completion allowance and rate limits. If the allowance runs out during reasoning, the final answer can be absent. A migration that preserves the old maximum-output setting can therefore change the rate of useful completions even when the request schema remains valid.

For the trial, record complete answers, accepted tool calls, retries, latency, and token consumption together. Start at the documented default and change effort against measured failures. Compare the total work needed to obtain an accepted result, including human repair. Lower generation compute does not settle that comparison if the surrounding agent needs additional attempts, loses necessary context, or requires a new parsing layer.

Capacity is another separate obligation. Reflection’s rate-limit documentation shares limits across an organization, counting input, output, and reasoning tokens. Adding projects or API keys does not increase that allowance. Daily limits reset at midnight UTC, and exhausted daily capacity is not repaired by immediate retries. These rules make a controlled trial possible to instrument, but they do not establish a production throughput commitment for an unspecified account.

The strongest case for trying Beam is straightforward: teams running text-only coding tasks through Chat Completions can investigate a new model without first owning its serving infrastructure. The strongest case against switching immediately is equally concrete: beta limits can change, the compatible surface is bounded, and the retrieved model and integration pages do not provide a paid per-token tariff. No dollar saving or self-hosted break-even point follows from these documents.

Today’s Mistral lead treats preview retirement policy as part of the purchase decision. Beam adds another requirement: the pilot must show that request budgets and endpoint behavior survive the move. Keep the current path available while collecting that evidence. Broader adoption becomes defensible when a confirmed commercial rate, stable quotas, released deployment artifacts, and lower measured cost per accepted task align. Until then, buy information with a bounded evaluation—not capacity on the strength of a compute chart.

Sources