skip to content
The Weighted Average

Models & Open Source

Arcee's 13B Active Model Still Has 400B Weights

Arcee's Series B strengthens its open-model case, but Trinity Large's 30.8x total-to-active parameter gap keeps self-hosting a separate decision.

Tweezers hold a small black microchip above a green circuit board
Tweezers hold a small black microchip above a green circuit board. Photograph by Vishnu Mohanan

Evaluate Trinity Large as a sparse large model, not a small model in disguise: Arcee’s Series B announcement values the company above $1 billion and backs another generation of open weights. Its existing flagship has 30.8x as many total parameters as active parameters per token, an architectural gap that keeps the self-hosting decision separate from the financing headline.

Sparse computation is not a small deployment

The financing statement identifies Trinity Large as a 400-billion-parameter mixture-of-experts model. The separate Arcee model overview lists 13 billion active parameters per token. Dividing 400 by 13 gives 30.7692, or 30.8x total parameters relative to the active count. That is an arithmetic description of the model, not a measured memory multiplier, throughput ratio, or comparison with another vendor’s system.

The distinction matters because the active figure describes which portion participates in a token’s computation; it does not say the remaining weights disappear from the deployment. Arcee’s own overview places Large in hosted endpoints or self-hosted multi-GPU configurations. A team sizing infrastructure should therefore ask for the actual artifact, precision, cache requirements, and serving layout rather than translate the active count directly into a device purchase. Sparse models can be efficient without being small enough for the machine already on the desk.

Arcee also reports approximately $20M spent building its entire 2025 model lineup, including salaries, compute, data, infrastructure, and operations. This is the company’s own accounting claim, not an independently audited model-training invoice. It is broader than the cost of one training run and broader than Trinity Large alone. Treating it as the price another organization could pay to reproduce the flagship would erase the scope Arcee explicitly describes.

The round’s amount is not disclosed in the announcement. Its stated valuation is not cash available to spend, and neither figure establishes a serving discount. What the announcement does provide is direction: next-generation Trinity models, expanded work with the Department of Energy and national laboratories, and products for customization, evaluation, deployment, and operation. That supports a supplier-continuity conversation. It does not replace a workload qualification.

There is already a concrete hosted alternative. Arcee’s pricing documentation lists Trinity Large Thinking at $0.25 per million input tokens, $0.80 per million output tokens, and $0.06 per million cached input tokens. These are token rates, not prices per completed task. A pilot should record actual input, output, cache use, retries, and accepted results before comparing that bill with a self-hosted configuration. The current prices make a controlled hosted evaluation possible without first committing to a multi-GPU deployment.

The product family offers different footprints rather than a single universal answer. Arcee’s Trinity page lists Nano at 6B total and 1B active, Mini at 26B and 3B, and Large at 400B and 13B. The page claims shared skills across sizes, but those marketing claims are not proof that prompts and accuracy remain interchangeable on a customer’s tasks. Smaller variants deserve their own tests, especially when local operation is the reason to consider them.

Make the endpoint prove the model card

The first test is less glamorous than a leaderboard: settle the endpoint contract. The model overview describes Large’s 512K model context as hosted at 128K, while the retrieved Trinity marketing page advertises a 256K context window. Those documents do not establish one unambiguous hosted limit. Ask Arcee to confirm the exact endpoint and deployed version, then test an input near the required boundary. Do not buy a long-context use case against whichever page happens to show the larger number.

This discrepancy need not mean the service is unreliable. Product pages and model documentation can describe different variants or update at different times. It does mean that a buyer lacks a sufficiently precise answer from those pages alone. Record the provider’s response in the acceptance criteria, alongside output limits, tool behavior, and the version identifier. Ambiguity resolved before migration is cheaper than ambiguity discovered when a production document is rejected.

The archive’s local DeepSeek inference analysis separates a kernel gain from end-to-end performance. Apply that discipline here without borrowing its hardware results. A favorable active-parameter count may reduce part of computation while cache pressure, communication, or output generation determines the actual delay. Measure full requests under the intended concurrency and context, not an isolated token rate on a convenient short prompt.

For builders whose main requirement is ownership, open weights still change the decision. They provide a path to operating and adapting the model rather than relying exclusively on a hosted endpoint. That path has responsibilities: packaging, evaluation, capacity planning, upgrades, and rollback. The funding announcement explicitly identifies operational tooling as an area of future investment, which is a useful admission that weights alone are not the whole product.

For builders whose main requirement is lower cost, begin with hosted evaluation unless a hard locality or control requirement rules it out. Compare accepted work and latency with the current system using the same tasks. Only then price the actual self-hosted configuration, including idle capacity and operational effort. The retrieved sources do not supply a measured self-hosted cost curve, so this article cannot honestly declare a break-even traffic level or a universal deployment winner.

Today’s MLPerf lead explains why reference-compliant performance is not application acceptance. Arcee deserves the same fair treatment: neither dismiss a sparse model because its total count is large nor approve it because the active count is small. A reproduced quality result, confirmed context contract, and measured serving bill would strengthen the switch case. Failed tool calls, unacceptable tail latency, or unresolved endpoint limits would weaken it. The Series B finances a roadmap; your workload still decides the purchase.

Sources