skip to content
The Weighted Average

Developer Tools

Kane CLI's Eight-Test Cap Is Not a Spending Cap

TestMu's new assurance workflow links requirements to tests, but an eight-test ceiling still creates 16 scenario/test records and uncapped agent work.

A person writing a checklist in an open notebook
A person writing a checklist in an open notebook. Photograph by Glenn Carstens-Peters

QA leads evaluating TestMu’s new Kane CLI Assurance Lifecycle should cap the test design without mistaking that cap for a spending limit. The documentation’s eight-pair example can still produce 16 scenario-and-test records, before acceptance criteria, gap records, or the work required to execute the tests.

A bounded suite can still have an unbounded bill

The September 16 announcement extends a browser-testing tool into requirements-based assurance. Instead of generating scenarios from a loose prompt, the workflow ingests requirement documents, extracts use-cases, designs tests, and connects execution evidence back to the requirements. The useful promise is traceability: an operator can ask which written requirement a passing test supports. That is a different claim from generating a larger pile of tests.

The derived count combines two pieces of official documentation. The automation guide demonstrates an unattended design invocation with --max 8. The design reference defines that ceiling as scenario-and-test pairs, with exactly one test per scenario. If the design fills the cap, eight scenarios plus eight tests make 16 records. That excludes the associated acceptance criteria and gap records, whose count depends on the actual requirements.

This is a structural count, not a measurement of reviewer hours or an estimate of production test coverage. Its importance is that the apparent unit on the command line is not the whole work product. A test may have several assertions; an acceptance criterion may need careful interpretation; an omitted scenario remains a risk decision. The design reference says evicted pairs become named budget gaps. A smaller retained suite does not mean the excluded obligations vanished.

Nor does --max cap agent turns or credits. The assurance overview identifies extraction, test design, and maintenance reconciliation as credit-consuming operations. Review, inspection, coverage, and store verification are local and free. The automation guide reports per-turn credits and cumulative credits as usage events. That gives teams a way to observe expenditure, but the documented scenario ceiling should not be sold internally as a dollar ceiling.

There is another cost after design. The coverage documentation says every newly designed test first needs an authoring run in a real browser. Before that, batch preflight reports missing recording metadata and the criteria remain unproven. Filling the eight-pair example therefore leaves eight tests needing their initial authoring pass. This follows from the documented per-test requirement; it is not a promise that those passes will succeed without retries or review.

The immediate buyer is a team that already has maintained requirements and needs an auditable connection to browser behavior. The wrong buyer is a team hoping the tool will infer a missing product specification and certify its own interpretation. Kane’s design documentation explicitly represents a missing expected result as a gap rather than an invented answer. That is a useful constraint only if somebody is responsible for resolving the gap.

Make the evidence survive the next branch and the next run

The storage boundary deserves attention before a pilot reaches CI. TestMu documents the local .context/ store as append-only, single-writer, and not git-mergeable; it warns that two branches appending records will corrupt the store on the next read. The recommended workflow is to keep it out of Git merges and share by re-ingesting sources. Teams accustomed to treating every generated artifact as a mergeable repository file need to change that assumption.

Local storage does not mean all processing is local. The overview says extraction and design agents run against the KaneAI service using the user’s login, while the store remains on disk. Review the requirement documents’ permitted processing boundary before ingesting sensitive material. The retrieved documentation establishes that architecture; it does not supply permission to upload every internal policy or customer specification merely because the resulting graph lives in the project directory.

Headless operation also needs an explicit policy. The automation guide distinguishes agent mode, which can pause on a high-risk question, from CI mode, which fails closed, and override mode, which accepts defaults. A resumable pause has exit code 3 and a 24-hour resume window. The same guide describes refusal and failure codes separately. A runner that treats every stopped process as success would erase the very assurance boundary the product is meant to preserve.

Choose the policy for the environment, not for the shortest green pipeline. An unattended release check should not quietly answer a high-risk ambiguity just to finish. A supervised agent can pause and ask for clarification. Store the status and the evidence together, and distinguish a completed design from completed execution. The tool’s output is useful precisely because those states are separate; flattening them into a single success badge throws away the information.

The coverage model is similarly deliberate. TestMu separates what an execution proved from what the live design still owes. It also joins tests to evidence through a hash of the resolved definition. A hand-edited test changes that identity and stops joining until the relevant redesign re-establishes the link. That can surprise teams used to manually patching generated tests, but it prevents an old passing result from silently certifying a different assertion.

Our Sourcegraph analysis found that outcome pricing leaves human review on the buyer’s tab. Kane makes the same operational point through a different product: generation, review, execution, and acceptance are separate costs. Measure them separately during a pilot. Existing deterministic tests may already cover a stable feature adequately; adding an agent and a new provenance store is justified only if the traceability or maintenance benefit exceeds the extra work.

Today’s Gemini Live lead examines another misleading shortcut from a visible unit to a complete bill. Here, the visible unit is a retained test pair, not a minute of audio. Start with one approved requirement set, a bounded design, and a documented release gate. Expand only when the team can reconcile credit use, human review, execution evidence, and remaining gaps. Eight tests can bound a suite; only measured work can bound its economics.

Sources