AI Economics for Operators
Harvey Trained Its Own Model for About $1.1M
Harvey post-trained Kimi K3 on roughly 150 B300s for two months — about $1.1M of rented compute — and claims SOTA on its own contracts benchmark.
Harvey, the legal-software company that built an eleven-figure business on other labs’ models, has now trained one of its own — and disclosed enough about the run to price it. In Harvey’s research preview of Tenet, the company says the model is a Kimi K3 base post-trained with Fireworks on “approximately 150 NVIDIA B300 GPUs over the course of 2 months.” At the $4.99 per GPU-hour floor that ComputeUnion’s B300 rental tracker records across seven providers on August 22, 2026, that run rents for roughly $1.1 million — 150 GPUs times about 1,464 hours in two months, times $4.99. Independent price aggregation from getdeploying’s B300 listings puts the market in the same band. Call it a mid-six-figure to low-seven-figure line item: less than one senior partner’s book of business, for a model that Harvey claims reaches first place on its own contracts benchmark.
That number is the story. Not because it is exact — reserved capacity, failed runs, data generation, and staff time all sit outside it — but because it is the first credibly disclosed order of magnitude for post-training a vertical model that competes with frontier systems on the tasks a specific industry pays for. For three years the build-versus-rent question in applied AI had one answer: rent, because training is a hyperscaler sport. Harvey’s disclosure moves the line.
The arithmetic that makes a wrapper stop being a wrapper
Start with what Harvey was paying before. Its platform routes lawyer requests across models from OpenAI, Anthropic, and Google, and Business Insider’s account of the Tenet launch frames the motive plainly: every call is a payment to a supplier that is simultaneously hiring its way into the legal market. Cofounder Gabe Pereyra told the publication that cost was one motivation and quality another. The base Harvey chose, Moonshot’s Kimi K3, lists at $3 per million input tokens and $15 per million output on Moonshot’s own pricing documentation, with cache hits at $0.30; OpenRouter’s Kimi K3 listing shows the same model at $2.60 and $13 across thirteen providers with a 1M-token context window. That is the input side of the trade: open weights with a competitive serving market underneath them.
The output side is where Harvey’s claims get interesting, because they are not primarily about benchmark scores. Post-training lifted held-out task completion on Harvey’s Legal Agent Bench to nearly double the base Kimi K3 rate, with a 9-point gain in all-pass rate; the company reports state-of-the-art on LAB Contracts and second place on LAB overall, using baselines drawn from the Vals LAB leaderboard. Fine. Every vendor’s benchmark flatters the vendor. The durable numbers are the efficiency ones.
On M&A diligence — tasks that require traversing up to 80 million tokens of data room — no baseline model or off-the-shelf coding agent passed more than 43.8% of rubric criteria. Moving to a recursive harness with a GLM-5.2 orchestrator reached 46.1%; post-training that orchestrator inside the harness via self-distillation reached 60.1%. The gain came from correcting a specific behavioral defect: the base model under-delegated review of the data room.
The agent harness, not the model, carried Harvey's biggest gain
Mean rubric criteria pass rate on LAB Diligence, 80M-token M&A data rooms
On firm-knowledge search, a Qwen3.8-27B model trained with Engram to internalize a firm’s corpus into parametric memory and structured notes cut tokens in completed trajectories by 58% and cost per query by 90%, tripling what Harvey calls intelligence-per-token — 190.8 rubric points per 100,000 inference tokens against 129.3 for the best frontier configuration, per Engram’s write-up of the memory work.
Read those three results together and the thesis is not “open weights caught up.” It is that the returns to post-training are concentrated in agent behavior, not knowledge — delegation, abstention, citation discipline, when to stop searching. Those are harness-shaped behaviors, and a frontier lab cannot train them for you because it does not have your harness. The same pattern runs through our earlier finding that naming an agent “coordinator” does nothing without trained coordination.
There is a second arithmetic worth running: the payback horizon. Harvey’s own document reports a Review Table model, trained with Applied Compute on a GLM-5.2 base, delivering answers at roughly one-tenth the cost per cell while improving citation quality by 12.1 points. A firm running 10,000-document reviews is the exact customer whose token bill scales linearly with matter size. If the $1.1 million of rented compute retires even a tenth of that recurring spend across a customer base of Harvey’s size, the run amortizes inside a fiscal year. That is the calculation every vertical software company should now be forced to perform out loud, because the inputs are finally public.
Who should copy this, and who should absolutely not
The temptation after a disclosure like this is for every Series B vertical-SaaS company to schedule a training run. Most should not. The preconditions Harvey met are expensive and specific.
First, a benchmark it owns. Harvey built LAB with more than 1,200 tasks across 24 practice areas, each with an expert rubric averaging 50 binary criteria, before it trained anything. Post-training without a graded environment is not training; it is expensive vibes. Second, expert data at scale: Harvey hired attorneys directly and through Mercor and Snorkel to author mock matters and grade rollouts. Third, a harness worth training inside — the diligence gain came from extending the agent loop with recursive sub-agents, then training the orchestrator against that specific loop. Fourth, the compute discipline to run group-sequence policy optimization with a rank-64 LoRA over a mixture-of-experts base for 150 optimizer steps and more than 10,000 rollouts per epoch without RL collapse.
The scoreboard results generalize better than the training recipe. Harvey reports that its gains transferred to benchmarks it never trained on, including Mercor’s APEX Agents corporate-lawyer leaderboard, where running base Kimi K3 in Harvey’s own harness lifted it from a reported 58.8% to 67.5% before any post-training at all. That single line deserves more attention than the model: an 8.7-point swing from harness choice alone, on an unchanged model. Most teams reading this have not exhausted their harness gains, and harness gains cost thousands of dollars, not a million.
The trade coverage read the launch through the vendor lens. Law.com’s report on Tenet notes it was tested against six rivals including Claude Fable 5, GPT-5.6 Sol, and Gemini 3.1 Pro. Artificial Lawyer’s assessment was blunter about what the absolute numbers show: on the hardest end-to-end agentic legal tasks, scores sit at “20% or lower” across the field. A model that leads a benchmark nobody passes is a leader in a race that has not started.
The ways this bet blows up
Four failure modes, in descending order of probability.
The benchmark is the product. Harvey grades Tenet on Harvey’s benchmark, with Harvey’s harness, judged by an LLM Harvey selected after ablations. The company is transparent about this — it documents where its runs diverge from Vals and Artificial Analysis methodology, and it had Mercor run APEX independently without disclosing task-level scores. But the structural problem stands: model makers write tests, then train toward them. Business Insider made the same point. Until a firm publishes matter-level outcomes on real files, the honest verdict is “promising internal results.”
The supply chain is geopolitical. Tenet’s base is a Chinese open-weight model. Harvey is selling to law firms whose clients include defense contractors, regulators, and sovereign-adjacent institutions. Weights are weights, and self-hosting removes the data path — but procurement committees do not reason about provenance the way engineers do, and the compute-control regime keeps tightening, as this week’s scrutiny of remote-access export loopholes shows. A single client mandate against Chinese-derived weights forces a re-base, and the training investment does not fully transfer.
Frontier labs collapse the gap by pricing. Harvey’s cost advantage rests on open-weight token prices roughly an order of magnitude below frontier list rates. That gap has compressed repeatedly this year — see Gemini 3.7 Flash’s price cut. The same vendors are also removing the procurement friction that made a specialist worth buying: Google has just folded its Antigravity coding agent into Gemini Enterprise seats rather than selling it separately. If a frontier vendor prices a legal-tuned tier aggressively while keeping quality, the million-dollar run becomes a rounding error against a supplier who can undercut you indefinitely.
The capital backdrop rewards the incumbents. Model suppliers are raising money at a scale no vertical can match, with Anthropic now targeting an IPO the size of the largest ever floated. Cheap capital buys the patience to price below cost in any vertical worth taking.
Nobody has shipped it. Tenet was not live in Harvey’s product at launch, and the company would not name testing firms or commit to a date. Every number above is a research preview.
What to do this quarter
The operator question is not “should we train a model.” It is “which of Harvey’s four preconditions do we already have, and what is the cheapest one that would move our numbers?”
- If you have no graded environment, build that first. A rubric-scored task set drawn from real work is the asset; the model is downstream. Harvey’s rubrics average 50 binary criteria per task and took months. Budget accordingly, and treat evaluation harnesses as the control plane rather than a reporting afterthought.
- Exhaust harness gains before compute gains. The 8.7-point APEX swing from harness choice alone cost nothing. Recursive delegation, forced abstention, and citation constraints are code changes.
- Price your own run before proposing it. Take your parameter count and target token budget, multiply GPU-hours by a live rental index rather than a list price, and add data generation — Harvey’s expert-authored corpus almost certainly exceeded its compute bill. If the total lands under one quarter of your model-vendor spend, the case is arguable; if not, keep renting and negotiate, as the teams reading token routing as a margin lever already do.
- Separate the compute bill from the data bill. Compute is the visible number and the smaller one. Harvey’s corpus required practicing attorneys authoring mock matters and grading rollouts; that labor market is itself now priced by vendors selling expert annotation, and it does not fall with GPU rental rates.
- Watch for the first customer-side result. The verdict changes when a named firm reports matter-level accuracy or realization-rate effects from Tenet in production. Until then, treat this as the most detailed public recipe yet for vertical post-training — not proof that the recipe pays.
Harvey’s valuation already assumed it would become more than a routing layer; the 44× revenue multiple attached to its $15.5 billion round priced a company that owns intelligence, not one that resells it. A $1.1 million compute bill is a cheap way to start defending that number. It is also, for everyone else, a published floor: the price of finding out whether your vertical’s hardest task responds to two months of training, rather than to a better prompt.
Sources
- Harvey — Tenet research preview, training and benchmark detail
- Harvey — Legal Agent Bench methodology
- Moonshot — Kimi K3 list pricing
- OpenRouter — Kimi K3 provider pricing and context window
- ComputeUnion — B300 hourly rental quotes across providers
- Engram — parametric memory results for legal agents
- Mercor — APEX Agents corporate-lawyer leaderboard
- Business Insider — Harvey’s motives for building Tenet in-house
- Law.com — Tenet benchmarked against six rival models
- Artificial Lawyer — how low the absolute agentic scores remain
- arXiv — group-sequence policy optimization, the RL method behind the run