Agentic Engineering
Docker's 61% Compute Discount Isn't a CI Replacement
Docker Cloud Sandboxes undercut a billed GitHub runner's hourly tariff by 61.1%, but lifecycle, policy and container limits decide whether agents should move.
Developers running unattended coding tasks should test Docker’s newly introduced Cloud Sandboxes, but keep their existing acceptance pipeline. The default $0.14 hourly compute tariff is 61.1% below the hourly equivalent of GitHub’s billed two-core Linux runner rate—a price comparison, not proof of cheaper completed work.
Cheap hours are not accepted changes
Docker’s proposition is straightforward: keep the microVM isolation used locally, run it on managed compute, and let the developer disconnect. The launch describes filesystem moves between laptop and cloud, preconfigured agent kits, network controls, and separately supplied model credentials. This targets long-running development work rather than another conversational interface. The useful question is whether an unattended task survives the move with its inputs, permissions, and result semantics intact.
The cross-vendor arithmetic needs its boundary stated immediately. Docker Small provides two vCPUs and 4 GiB for $0.14 per hour. GitHub lists its standard Linux two-core x64 runner at $0.006 per minute, or $0.36 for a fully billed hour. Calculate one minus $0.14 divided by $0.36: 61.1%. The nominal core counts align; hardware, memory, platform features, scheduling, and task throughput need not. This is a tariff screen for a pilot, not an equivalent-machine benchmark.
The saving also disappears as a billing argument when the competing minute is free. GitHub’s billing documentation makes standard public-repository runners free and gives private accounts included usage. A team below its included allowance does not save $0.22 by moving an hour elsewhere. It adds a new cash cost unless the change buys something else, such as a different lifecycle or less dependence on a laptop. Use the marginal bill your team actually faces.
Docker’s price curve contains another useful finding. The launch lists hourly rates of $0.07, $0.14, $0.28, $0.56, and $1.12 for Micro through XL. Divide those by their respective one, two, four, eight, and sixteen vCPUs: every size costs $0.07 per vCPU-hour. Memory scales alongside compute at 2 GiB per vCPU. Moving up a tier buys resources, not a published per-core discount. The larger machine earns its keep only if it finishes useful work faster or enables a task the smaller one cannot run.
Docker's bigger sandboxes offer no per-core discount
USD per vCPU-hour, hourly price divided by vCPUs; inference billed separately
Inference remains a separate bill. Docker’s subscription documentation says compute is metered by the second and model providers charge independently. The launch’s temporary credits are therefore a way to test the infrastructure, not evidence of sustainable agent economics. The experiment should record compute, model usage, retries, and review effort separately. Otherwise a cheap host can hide an expensive loop that produces no acceptable change.
Our Strands analysis used cost per successful benchmark task rather than cost per attempt. Apply the same discipline here: define acceptance before comparing bills. A pull request, an agent’s completion message, and a passing check in the correct environment are different outcomes. Keeping that distinction is the cheapest protection against buying apparent productivity.
A move copies files, not your operating assumptions
The headline command sounds like a handoff, but the move documentation describes a filesystem snapshot and a separate destination sandbox. Running processes and memory do not transfer. The source remains in place. That makes the feature useful for relocating a prepared environment, not for assuming an in-flight process migrates transparently. Plan a checkpoint that can be restarted and verified, rather than treating a move as live process continuity.
The local-versus-cloud comparison lists separate resource stores: secrets, network policy, templates, and MCP configuration do not become universal merely because the CLI is familiar. Host mounts and clone-mode volumes are not copied by the move, and cloud sandboxes cannot reach back into local paths. A repository present through a host integration needs an explicit transfer or clone strategy. File presence is an acceptance check, not an assumption.
This matters especially for the instructions that shape agent behavior. Our AGENTS.md portability coverage distinguished shared project guidance from tool-specific override behavior. A similar test belongs at the environment boundary: verify that the destination sees the intended repository, revision, and project instructions. Do not debug an apparent reasoning regression before checking whether the agent received the same working context.
Credentials need their own migration plan. Docker documents distinct local, cloud-CLI, and Agentic Platform secret interfaces. A credential created through one is not automatically available through the others. Managed cloud credentials stay outside the sandbox filesystem, while interactive agent sign-in can write credentials into files that snapshots may carry. Teams should use the supported managed store and inspect what a snapshot will contain before making it a reusable template.
Network policy has an equally important boundary. Cloud rules control destinations but do not support HTTP method or path restrictions. Local policies are not copied. If a local workflow relies on narrower controls, a successful filesystem transfer does not establish equivalent authorization in the cloud. Security owners should verify allowed and denied connections in the destination and decide whether the available controls satisfy the task before unattended execution begins.
The commercial identity is also relevant. Docker’s signup page describes an experimental Agentic Platform subscription attached to a personal Docker account and billed separately from the Docker subscription. Meanwhile, the sandbox FAQ distinguishes separately paid organization governance for local environments. A team should not infer enterprise cloud governance from a developer’s ability to start a sandbox. Procurement needs the actual control scope, account ownership, and invoice path.
None of this makes the product incoherent. Separate stores can be a sensible way to avoid silently exporting local privileges. It does make the migration a small integration project rather than a cost-free flag change. The first pilot should produce a short, reproducible destination checklist so the next developer does not have to rediscover which assumptions stayed on the laptop.
The documented limitation that can falsify a test
The most consequential caveat is not price. Docker’s cloud usage guide documents a container execution and healthcheck limitation: an exec operation can access the sandbox VM filesystem rather than the intended container filesystem. Commands can fail, or succeed while touching the wrong files. Compose healthchecks share the affected path. An apparently successful exit status is therefore insufficient evidence that the intended container was tested.
For teams whose acceptance depends on container-local setup, debugging, or readiness checks, this should be a migration blocker until the exact workflow is demonstrated to work. The documentation says a container’s main process uses the correct filesystem, which may support a narrower test arrangement. It also cautions that this does not restore general exec behavior or ongoing health monitoring. Rewriting an entire test system around a temporary workaround could erase a small compute saving.
Lifecycle semantics create a second boundary. The usage guide says cloud sandboxes expire after one hour by default, and expiration extensions cannot push beyond 24 hours from creation. Depending on support and the chosen timeout action, expiration stops or deletes the sandbox. A developer planning overnight work must set and verify the intended outcome before leaving. An always-available cloud service does not mean an individual sandbox has an unlimited lifetime.
Persistent volumes require particular care. Docker describes cloud volumes as experimental, with data saved when a sandbox exits rather than continuously. If multiple sandboxes mount the same volume, the last one to exit overwrites the stored snapshot. That is not a shared transactional workspace. Parallel tasks should have independent state and explicit result collection; otherwise the very concurrency the product enables can make outputs harder to trust.
The strongest case against moving is consequently boring: the existing system may already be good enough. A team with free runner minutes, working credentials, reliable artifacts, and a stable acceptance pipeline gains little from chasing a lower published hourly number. A team whose agents are repeatedly interrupted by local availability has a different problem and a stronger reason to test managed execution. Segment the users before announcing a platform-wide migration.
Evidence that would change the recommendation is equally concrete. Reproduced task-level savings, intact destination policies, dependable output retention, and correct container execution would justify expansion. Unexplained discrepancies between environments, review work that outweighs tariff savings, or unsupported controls would justify keeping execution where it is. This is not a benchmark result reported by the paper; it is the evaluation required before the pricing comparison becomes a purchasing conclusion.
Move the experiment before you move the pipeline
Start with a bounded development task that can be restarted and whose result can be checked elsewhere. Give it a fixed input revision and explicit acceptance criteria. Record the selected size, elapsed billed compute, model charges, output artifacts, and reviewer decision. Run the existing acceptance checks in the established environment. That design tests the new execution surface without asking it to certify itself during the same experiment.
Separate resource tuning from agent tuning. If a larger sandbox finishes sooner, measure the cost of the accepted result rather than celebrating wall-clock speed alone. Because the published per-core price is constant, extra cores are not inherently a bargain. If the task remains bottlenecked on model calls or human review, a bigger machine may add expense without resolving the constraint. Let the workload demonstrate which resource it needs.
The same evidence boundary appears at larger scales in this edition. Anthropic’s Akamai commitment changes an infrastructure investment burden, not a public unit tariff. Boom’s lost Crusoe order separates announced power demand from an accepted supply plan. Docker is unusually testable by comparison: the unit price is public and the experiment can be small. That is a reason to measure, not a reason to skip measurement.
Before wider use, assign an owner for account billing, expiration behavior, network policy, credentials, and artifact recovery. Those responsibilities exist even if no dedicated orchestration service is needed. Avoid building one until repeated failures demonstrate the need. A clear task contract and retained acceptance evidence will tell the team far more than a dashboard counting how many agents were launched.
The immediate verdict is selective adoption. Developers with restartable, independent tasks and a genuine need for remote continuity should pilot Cloud Sandboxes. Teams depending on unsupported local controls or affected container operations should delay the move. The public tariff clears an economic screening test; it does not clear operational acceptance.
- Platform engineers: pilot one restartable task class, retain the existing acceptance pipeline, and verify that destination files and project instructions match the intended revision.
- Finance owners: compare marginal paid compute, not free included minutes; add model charges and review effort before calling the 61.1% tariff gap a saving.
- Security owners: configure cloud credentials and destination policy explicitly. Reject tasks whose required control boundary the cloud environment cannot reproduce.
- Reliability owners: test expiration, result recovery, and container behavior before unattended use. Expand only when a successful result means the same thing on both sides.
Cheap compute is useful only when the execution boundary preserves the meaning of the result.
Sources
- Docker — Cloud Sandboxes launch, resource sizes and hourly prices
- GitHub — Actions runner minute tariffs
- GitHub — free and included Actions usage
- Docker — Agentic Platform subscription and separate inference billing
- Docker — filesystem move semantics
- Docker — local and cloud resource differences
- Docker — cloud credentials and snapshot boundaries
- Docker — cloud destination-based network policy
- Docker — sandbox FAQ and local organization governance
- Docker — container exec and healthcheck limitation