skip to content
The Weighted Average

Agentic Engineering

Strands Harness Cuts Benchmark Cost per Pass to $0.91

AWS's Strands harness implies $0.91 per passed Terminal-Bench task versus $4.51 for Claude Code. Reproduce the vendor test before migrating.

Two railway tracks diverge through a green forest
Two railway tracks diverge through a green forest. Photograph by Sergej *****

Strands has released an open general-purpose agent harness claiming 28% lower token cost across six matched-model benchmarks. Its detailed Terminal-Bench comparison, combined with the benchmark authors’ 89-task specification, implies $0.91 per passed task for Strands against $4.51 for Claude Code—a vendor-reported result worth reproducing before moving production agents.

Divide the bill by work that passed

The launch’s embedded benchmark publishes the underlying costs and scores: Strands with Claude Fable 5 costs $56.29 at 69.66% accuracy, while Claude Code with the same model costs $248.05 at 61.80%. The embed says each harness ran 89 trials. The independent Terminal-Bench release identifies 89 tasks in the suite, providing the dataset boundary for the calculation rather than leaving the denominator implicit.

Normalize spending by the number of passes implied by each reported rate. Strands gives $56.29 ÷ (89 × 0.6966) = $0.91; Claude Code gives $248.05 ÷ (89 × 0.6180) = $4.51, rounded to cents. The result is an approximately 4.97-fold gap in token cost per implied passed task. Because published scores are rounded, these calculations do not reconstruct exact integer success counts or unpublished trial records.

Strands' vendor test implies roughly 5x cheaper passes

Token dollars per implied pass. Same Fable 5 model, 89 trials; excludes hosting and review.

Claude CodeStrands$0$1$2$3$4$5$4.51$0.91
Claude CodeStrands$0$1$2$3$4$5$4.51$0.91
Strands vendor benchmark; Terminal-Bench 2.1 task count · Sep 22, 2026

This is stronger evidence than a cheaper aggregate run by itself. A harness could spend less by stopping early and failing more often. Here the vendor reports both lower spending and a higher pass rate for Strands in the selected comparison. It does not follow that the same advantage holds for every model, benchmark, or business workflow. The launch’s broader 28% claim and this particular cost-per-pass result describe different aggregation levels.

The mechanism is concrete. Strands says tool results above roughly 1,500 tokens are truncated and compaction triggers above 85% of context capacity. It also describes prompt caching and context recovery after overflow. These defaults change what the model repeatedly reads. The hypothesis is not that identical requests become magically cheaper; it is that better management of the agent’s working context can reduce unnecessary token consumption without sacrificing useful evidence.

That last condition deserves testing. Truncating a long tool response can remove noise, but a task may need information deep in that response. The harness documentation describes setting bulky results aside and retaining sessions and memory. Verify that your tools and prompts recover the needed material after offloading. A savings policy is only helpful if the agent can still find the evidence required to finish correctly.

The artifact is deployable rather than merely a benchmark scaffold. The Apache-2.0 SDK repository supports Python and TypeScript, and the launch describes local use or deployment in Linux containers across providers. Teams building their own agents can therefore test the runtime separately from a change in model vendor. Open licensing lowers one adoption barrier; it does not eliminate hosting, inference, observability, or maintenance costs.

A harness swap needs a regression budget

The first candidates are teams whose existing model is adequate but whose agent repeatedly drags large histories and tool outputs through long tasks. A bounded harness comparison can isolate that problem more cleanly than changing the model, prompts, tools, and deployment simultaneously. Keep the model and acceptance rules fixed initially, then inspect why token use changes. Otherwise an apparent efficiency gain can hide a weaker task definition.

Pin the benchmark version too. The Terminal-Bench authors say version 2.1 fixes 28 of version 2.0’s 89 tasks, including external-dependency drift, resource mismatches, and instructions that did not match tests. A score from the older suite is not a clean historical baseline for this release. The revision demonstrates why benchmark names alone are insufficient provenance: task environments and evaluation rules can change the result without any model improvement.

There is an important cost caveat in the vendor’s own embedded data. It reports about 28.7 million tokens for Strands and 55.0 million for Claude Code, while the dollar-cost ratio is larger than that token-count ratio. Aggregate tokens do not explain the bill on their own; input, output, cache treatment, and exact pricing configuration matter. The fetched launch does not provide a complete per-token accounting reconciliation. Request that detail before treating the monetary gap as portable to your contract.

Nor should the benchmark bill be mistaken for a production total. Strands’ article describes testing on EC2 with Harbor, but the headline concerns token cost. Your deployment still needs runtime capacity, logs, evaluation, and an owner for failures. Our earlier AgentCore analysis separated evaluation spending from application inference. A cheaper generation loop does not establish that the surrounding assurance system became cheaper too.

Today’s Grok 4.7 lead makes the complementary point about tariff configuration. Model rates and harness behavior are separate levers. Test them independently before combining claimed savings; multiplying a vendor’s harness discount by another vendor’s price advantage would assume compatible workloads and quality that neither source establishes.

For the pilot, preserve accepted outputs, failed attempts, exact model settings, token categories, elapsed time, and human interventions. Include tasks with large files and long tool results, since those are precisely where the default context policy should matter. Compare recovered evidence and final correctness, not merely shorter traces. Also test session continuation and restart behavior if the intended deployment depends on durable work.

The verdict is a controlled runtime evaluation, not a mandate to abandon an existing coding agent. Evidence that would justify migration is a repeatable drop in cost per accepted result without losing tool coverage, reviewability, or reliability. Evidence that would reverse it is a saving driven by missing context, incompatible semantics, or accounting assumptions that disappear under your provider terms. Strands has supplied a substantial hypothesis and inspectable implementation. The remaining purchase is the work required to prove it on your tasks.

Sources