Enterprise AI & Work
Salesforce Koa's Open Base Is Not a Self-Hosting Offer
Salesforce's Koa enters pilots on a 120B-parameter base. Its 10:1 total-to-active parameter ratio does not make the hosted model a small local download.
Salesforce administrators should evaluate Koa’s new Agentforce pilot as a hosted CRM specialization, not a self-hosting shortcut. Its underlying model has a 10:1 ratio of total to active parameters—a useful warning that the attractive active-parameter count is neither the full model footprint nor a customer price.
Open weights give Salesforce control, not every customer a server
Salesforce and Nvidia announced Koa on September 15. Salesforce says it post-trained Nemotron 3 Super for enterprise workflows and operates the resulting model inside its own trust boundary. That is a meaningful deployment distinction for a company already using Agentforce. It means the announcement’s control story concerns where Salesforce runs inference and training, rather than an offer to download Koa into any customer’s infrastructure.
The technical lineage makes the sizing trap visible. Salesforce’s paper identifies the foundation as Nemotron-3-Super-120B. Nvidia’s corresponding model documentation specifies 12 billion active parameters out of 120 billion total. Combining those records gives 120 ÷ 12 = 10:1 . That is an architecture ratio for Koa’s identified base, not a measured serving footprint for Salesforce’s post-trained deployment.
A model does not become a small dense model because only part of its parameter set is active for a token. Buyers considering the open foundation separately must budget the full serving configuration, including weights, runtime state, and concurrency. The ratio alone does not supply a memory figure, a throughput estimate, or a tenfold discount. Salesforce’s hosted service also does not pass through an infrastructure saving merely because its supplier publishes an efficient architecture.
Our PAIR analysis distinguished request routing from pooled model memory. Koa creates a related boundary: an open foundation gives the service builder architectural options, but those options are not automatically product entitlements for the service buyer. A team should distinguish using hosted Koa, deploying Nemotron itself, and commissioning a separate specialized model. They require different work and produce different support obligations.
Availability reinforces that distinction. The launch says Koa is available to selected pilot customers, with general availability expected in Winter 2026 in U.S. regions. Separately, Salesforce describes post-trained Nvidia models for Missionforce Operations, available to selected customers in October, including specialized deployment environments. The announcement does not make those timelines or deployment rights interchangeable. A Koa pilot is not proof that the customer can deploy the same model air-gapped next week.
Training claims also merit precise reading. The launch describes synthetic CRM scenarios and says no customer data was used. The technical paper describes public and synthetically generated training data, likewise excluding customer data. Those statements support a no-customer-training-data claim in the reported development process. They do not establish every contractual rule governing future customer inputs, retention, or administrative access. Procurement still needs the terms for the actual pilot.
The benchmark is promising; the purchase is still conditional
The full Koa paper supplies a more restrained picture than the launch superlative. On its CRM benchmark, Koa has overall accuracy of 0.86 against its base model’s 0.84, and function-call accuracy of 0.77 against 0.71. The authors say it improves on the open base while remaining below the strongest frontier models across the broader evaluation. That supports specialization, not a conclusion that every enterprise should replace its current model.
The error arithmetic illustrates why benchmark labels matter. For the paper’s function-call comparison, residual error falls from 0.29 to 0.23. The relative reduction is (0.29 − 0.23) ÷ 0.29, or 20.7%. This is not automatically the same comparison as the launch’s “three times fewer errors” claim: that headline does not identify the identical baseline and slice in the announcement text. Ask the vendor to reconcile the exact workload, comparator, and scoring rule before using either figure in a business case.
A pilot should reproduce the operations the customer actually needs: the relevant CRM objects, approval conditions, and sequences of tool calls. Preserve both the intended result and the final stored state. A plausible response is not enough when the job is to update an opportunity correctly. This is a proposed acceptance process, not an assertion that the public benchmark has measured the buyer’s configuration or organizational policies.
The cost question remains open. Neither the launch nor the research paper publishes a Koa-specific customer tariff. It would therefore be misleading to translate active parameters or benchmark gains into dollars saved. Request the billing unit, any model-specific premium, the treatment of retries, and the existing Agentforce entitlements required to participate. Include implementation and review work in the comparison rather than treating those costs as a sunk platform expense.
The strongest case for adopting is concentrated expertise within an existing Salesforce workflow. If a team can improve accepted CRM actions without introducing another inference boundary or rebuilding its integration, specialization may be more valuable than a higher score on an unrelated general benchmark. The strongest countercase is that the existing configuration already performs adequately, or that the proposed pilot lacks the regional availability and contractual terms the organization needs. Either can justify waiting.
Today’s Gemini Live lead distinguishes fluent conversation from completed work. Koa applies that test behind the conversation: did the right action occur under the right policy, and what did it cost? Evidence that would change the verdict is a representative customer evaluation with fewer consequential errors, transparent action-level costs, and confirmed deployment terms. Pilot the hosted specialization where it fits; do not buy a self-hosting narrative the product has not offered.