skip to content
The Weighted Average

Agentic Engineering

AWS AgentCore Sampling Can Add $10,692 a Month

At AWS's example workload, moving evaluation sampling from 2% to 20% adds $10,692 monthly. Budget quality checks separately from incident response.

An airport departure board showing flight numbers and gate-closing notices
An airport departure board showing flight numbers and gate-closing notices. Photograph by CHUTTERSNAP

Platform teams adopting AWS’s dual-layer agent monitoring walkthrough should budget continuous quality evaluation separately from incident investigation. Applying the different sampling settings in AWS’s own examples to one published workload turns a $1,188 monthly evaluation bill into $11,880—a $10,692 difference, not an observed customer invoice or a reason to stop checking quality.

The green dashboard misses the failed booking

AWS’s walkthrough pairs AgentCore Evaluations with AWS DevOps Agent around an airline-reservation demonstration. One layer scores whether agents helped users accomplish their goals; the other investigates infrastructure problems. The distinction is commercially important. A service can return a response while choosing the wrong specialist, misunderstanding a request, or failing to finish a booking. Infrastructure health and task completion are related signals, not interchangeable ones.

The demonstration uses a supervisor and specialist agents for flights, users, and reservations, with dynamic handoffs rather than a fixed execution path. AWS describes how missing permissions and other infrastructure problems can surface as degraded behavior instead of a clear application error. Its examples illustrate a failure class; they are not evidence that a measured percentage of production airline requests failed in the wild. The reader should adopt the distinction without inheriting an invented incident rate.

The AgentCore evaluation overview describes integration through OpenTelemetry and OpenInference instrumentation, including Strands and LangGraph. That makes existing traces useful inputs to quality assessment rather than forcing a team to infer success from raw infrastructure counters. However, converting traces into scores adds model work. The score is not a free annotation attached to an already-paid request.

There is no need to move an otherwise suitable application merely to copy the demonstration’s hosting choice. AWS’s service mechanics documentation says evaluation can assess agents hosted inside or outside AgentCore Runtime. Teams should first establish whether their traces contain the evidence needed to judge the task. Runtime migration is a separate decision, with its own costs and operational consequences.

This extends the archive’s distinction between valid tool calls and useful business outcomes. Monitoring needs to observe the property the operator cares about. For a reservation workflow, that may mean a valid booking under the required policy, not a fluent summary of available flights. Define that result before purchasing more dashboards; otherwise the new metrics can reproduce the old blindness with better typography.

Three examples, three very different bills

The cost model is unusually easy to reconstruct because AWS’s AgentCore pricing page publishes a worked customer-service example. It evaluates 10,000 production conversations monthly at a 2% sample, implying 500,000 production interactions. Each built-in evaluation consumes 15,000 input tokens and 300 output tokens. The quoted rates are $2.40 per million input tokens and $12 per million output tokens, with the underlying model cost included.

That gives 15,000 × 2.40 ÷ 1,000,000 + 300 × 12 ÷ 1,000,000 = $0.0396 per judge evaluation. The walkthrough selects 3 built-in evaluators—helpfulness, correctness, and goal success rate. Following the pricing example’s convention of one evaluation for each judge per selected interaction gives $0.1188 per evaluated interaction. Real traces and sessions can differ; this is the published example’s accounting unit, not a universal mapping from a user request to three charges.

Now change only the sample. The walkthrough describes 10% as a typical production sampling rate, while the companion sample repository’s configuration uses 20%. Applied to the same inferred 500,000 monthly interactions, the three published settings select 10,000, 50,000, or 100,000 interactions. Multiply each by $0.1188 and the monthly built-in evaluation costs are $1,188, $5,940, and $11,880 respectively.

The largest difference is $11,880 − $1,188 = $10,692 per month. Every input comes from the pricing example, walkthrough, or sample configuration. This combines their assumptions to expose a configuration consequence; it does not claim that AWS recommends the same sample for every workload or that its demonstration incurred these charges. Higher sampling buys more observations, but the example alone cannot tell a buyer how much additional harm those observations prevent.

Sampling choices turn $1,188 into $11,880 a month

Same 500,000 monthly interactions and three judges · illustrative, not measured

2% sample10% sample20% sample$0$5K$10K$11,880$5,940$1,18810× the 2% bill
2% sample10% sample20% sample$0$5K$10K$11,880$5,940$1,18810× the 2% bill
AWS AgentCore pricing, monitoring walkthrough, and sample repository · September 2026

The comparison deliberately excludes the pricing page’s fixed development test workload and custom evaluator. It also excludes application inference, hosting, telemetry, and incident investigations. Keeping those items outside all three bars makes the sample comparison like-for-like. They still belong in a production budget. In particular, CloudWatch prices ingestion, storage, and analysis separately; cheaper evaluation does not erase the cost of retaining and examining traces.

Input length is the quieter multiplier. In the same pricing example, input accounts for $0.036 of each $0.0396 judge evaluation, or 90.9% of the token charge. That follows directly from the published token counts and rates, not from a measured production distribution. The example includes conversation history and business context in its input. A team that retains the sampling percentage while sending longer histories can therefore change the bill substantially. Measure evaluator input separately from the application’s final answer length.

Incident response has a different meter. The AWS DevOps Agent pricing examples charge $0.0083 per active agent-second and illustrate a small team’s ten eight-minute investigations at $39.84 monthly. That is a separate illustrative workload, not a prediction for the airline system. Combining its small bill with continuous evaluation without preserving the two denominators would hide the main budgeting choice: how much routine traffic gets scored, not merely how often an investigator runs.

A cheaper sample can still be the wrong sample

The tempting response is to select the smallest percentage and declare victory. That would confuse cost control with adequate coverage. AWS’s online evaluation documentation defines configuration around evaluators, data sources, and evaluation parameters. A production owner should choose those deliberately. The right sample depends on which failures matter, how often the relevant workflow occurs, and what other checks already protect the result—not which setting happened to appear first in a tutorial.

A low-volume, consequential workflow can disappear inside an aggregate score dominated by routine requests. That is a reason to inspect coverage by task type and review known failures separately, not a claim about a measured failure distribution in AWS’s demonstration. Keep complaint-triggered and regression evaluations alongside broad sampling. If the expensive errors do not enter the evaluation set, a cheaper bill and an attractive average may coexist with a worse service.

Judge reliability is the second constraint. AWS’s walkthrough explicitly warns that model-based scoring lacks ground truth and recommends calibration with domain experts. The built-in evaluator documentation also says its models and prompt templates cannot be modified. A convenient default may therefore be a poor fit for a specialized acceptance rule. Test disagreements against human-reviewed examples before wiring a score directly into consequential deployment decisions.

Custom evaluation is an alternative, not an automatic discount. AWS’s pricing page lists $1.50 per thousand custom evaluations, but explicitly excludes the customer’s additional model-inference cost from that platform fee. Comparing it with a built-in rate that includes model usage would mix price boundaries. A custom judge may be worthwhile because its rubric better reflects the task; establish that benefit first, then compare the complete cost. Replacing a well-calibrated default merely because its displayed fee looks larger is the wrong optimization.

The timing boundary matters just as much. The walkthrough describes online evaluation as asynchronous and sampled: a problematic response can reach the user before it receives a score. Bedrock Guardrails offers configurable content, topic, grounding, and sensitive-information checks as a complementary layer. Neither product description establishes that a booking is authorized or a financial action is appropriate. Business permissions, validation, and approval controls still need to operate where the action is committed.

That is the same boundary explored in our Docusign analysis of access versus approval. Permission to invoke a tool is weaker than permission to perform every action it exposes. An evaluator that later dislikes an action cannot undo an irreversible consequence. Keep preventive controls separate from retrospective measurement even when both appear in the same monitoring console.

The strongest argument for spending more is that quality failures may cost far more than their detection. The published examples do not quantify that avoided loss, so this article does not manufacture a return-on-investment estimate. Evidence that would justify the higher sample is additional actionable detection, faster validated recovery, or demonstrably better coverage of important failures. Evidence against it is a rising judge bill that mostly repeats low-value observations without changing a decision.

Make the monitoring budget earn its place

Start with a narrow workflow whose desired result can be checked. Preserve a small set of accepted outcomes, known failures, and difficult boundary cases. Confirm that the traces contain enough evidence to distinguish them. Then compare evaluator judgments with that record before expanding coverage. Buying a larger sample before establishing judge usefulness merely scales uncertainty and expense together.

Keep the unit of accounting explicit in the implementation. Sessions, traces, tool spans, judge invocations, and user interactions need not have identical counts. The pricing example simplifies that mapping for arithmetic; a real deployment must measure it. Record selected interactions, actual judge token usage, and evaluators invoked, then reconcile those observations with the invoice. A sampling percentage alone cannot explain costs if traces grow longer or additional evaluators run.

Keep offline release testing separate from live monitoring. AWS’s pricing page offers a batch-evaluation discount, while the walkthrough uses on-demand checks for specific sessions and continuous sampling for production. Those modes serve different clocks. A release regression can wait for an offline result; an emerging production problem needs timely observation. Decide which questions each mode must answer and budget them independently. A lower batch rate is useful for suitable work, not a reason to replace live coverage with yesterday’s test set.

The same discipline applies when generation becomes cheaper. Today’s Fugu Max brief examines the gap between token rates and accepted-task costs. If a cheaper model needs more review, the evaluation layer may consume part of the saving. That can still be a good trade, but only when both sides are measured. Do not optimize generation spend while treating verification as an unbounded shared service.

Dependency changes also belong in the quality loop. DeepSeek’s conflicting Pro migration notices show why a familiar identifier is not sufficient evidence of continuity. A provider change should trigger targeted regression checks even when infrastructure remains healthy. Keep the event that prompted evaluation alongside the resulting scores so an operator can distinguish model changes, prompt changes, tool failures, and ordinary traffic variation.

The operating model is therefore a budget with an owner, not a universal sample recommendation. Finance should see the quality bill separately from application inference and incident response. Engineering should own the mapping from selected traffic to billable evaluation work. The service owner should decide what evidence merits a larger sample. When nobody owns all three views, a copied configuration can become a recurring expense without an explicit decision to buy it.

  • Agent platform teams: adopt outcome evaluation when infrastructure metrics miss task failures; start with calibrated judges and a trace audit rather than a wholesale runtime migration.
  • Engineering and finance leads: budget actual judge tokens, evaluator count, sampling, telemetry, and incident investigations separately. On AWS’s example workload, 20% rather than 2% sampling adds $10,692 monthly before the excluded services.
  • Workflow owners: retain synchronous authorization and validation for consequential actions. Expand sampling when measured detection or coverage improves, and revisit it when model behavior, trace length, or task mix changes.

The verdict is to buy both kinds of visibility without confusing their economics. AWS’s architecture provides a useful separation between agent quality and infrastructure diagnosis. Its own examples also show why the sampling knob deserves a budget review. A green dashboard is not a receipt for completed work; a larger evaluation invoice is not one either.

Sources