Agentic Engineering
Atlas Infinite's 128 TB Preview Lacks Vector Search
MongoDB's new preview separates storage and compute, but its 128 TB logical ceiling does not deliver Vector Search, live migration, or an uptime SLA.
Agent teams should separate a database upgrade from an architecture migration after MongoDB’s September 29 launch of MongoDB 9.0 and Atlas Infinite. Infinite’s 128 TB logical-storage ceiling is nominally 32x the standard M60 configured-storage ceiling, but the preview does not support Atlas Vector Search, cross-edition live migration, or an uptime SLA.
The bigger container is not the whole agent stack
The opening ratio joins two documents and needs its units intact. Infinite’s overview permits up to 128 TB of logical storage per replica set. The AWS configuration table lists standard M60 storage up to 4 TB. Dividing 128 by 4 yields 32x, a ratio of published ceilings. Logical data and configured disk capacity are different measures, so this is not a claim of 32 times the usable application data, throughput, or value. Extended-storage configurations are also outside that standard-tier comparison.
The meaningful change is architectural. Infinite separates the compute layer from managed storage, so resizing compute need not move or replicate the underlying dataset. That addresses a genuine purchasing question: can a storage-heavy workload avoid buying more compute merely to hold more data? It does not answer whether an existing agent application can move intact, because compatibility includes services and operations surrounding the database API.
MongoDB launched three related products with different readiness states. MongoDB 9.0 is generally available; Infinite and Atlas Agent Engine are public previews. Agent Engine combines runtime, memory, retrieval, and governance, and lets customers adopt components independently. Its consumption charges can draw on existing Atlas commitments. None of that establishes that every feature in the Agent Engine announcement is supported by an Infinite cluster today.
The Infinite limitations are decisive for retrieval applications. Atlas Search and Vector Search are absent from the preview, as are sharded, global, multi-region, and multi-cloud clusters. Availability is AWS-only, within a single region. An agent team depending on those services should preserve its supported deployment rather than read the shared launch date as a compatibility matrix. A modular platform can be useful without every module working on every new storage architecture.
That leaves a narrower, credible audience: teams with compatible transactional workloads, growing retained data, and a reason to test independent storage and compute scaling. Their first move should be a separate evaluation cluster, not a production cutover. Teams whose immediate bottleneck is retrieval quality or agent governance can evaluate those layers without assuming a database-edition migration is required.
Our MLPerf analysis distinguished pipeline throughput from accepted answers. MongoDB introduces another boundary beneath that one. A larger database ceiling does not certify retrieval support, and retrieval support does not certify answer quality. Keeping these tests separate prevents the infrastructure business case from inheriting capabilities that were announced elsewhere in the same presentation.
Storage decouples; the performance pipe does not
Independent capacity does not mean independent throughput. MongoDB’s AWS specifications publish storage throughput of 80 MBps for M30, 150 for M40, 300 for M50, and 600 for M60 General Infinite clusters. The top value divided by the first is 7.5x. These are tier specifications, not measurements of an application. They make the operating constraint visible: a large logical dataset still passes through a tier-bound storage interface.
Atlas Infinite's M60 pipe is 7.5x its M30 pipe
Published storage throughput by General cluster tier, MB/s · September 2026
The autoscaling documentation explicitly says storage IOPS and throughput change with the compute tier. Those values cannot be adjusted independently within a tier. A buyer may therefore gain freedom from storage-capacity-driven overprovisioning while still needing a larger compute tier for storage performance. Evaluate working-set access and query behavior; do not assume the cheapest supported compute tier can efficiently serve the largest permitted dataset.
MongoDB’s launch reports 189% more throughput per dollar than Atlas Core in internal testing. That means 2.89 times the throughput per dollar under the tested conditions, not a 189% reduction in cost. At identical required throughput, its reciprocal implies about 34.6% of the reference spending, or a 65.4% reduction, if the test’s relationship carries over. This is conditional arithmetic, not a quoted customer discount or an independently replicated operating result.
The Infinite billing guide specifies separate compute, storage, backup, and data-transfer charges. Compute is billed for provisioned nodes even while idle. Storage is metered as uncompressed logical data in gigabyte-hours, including the oplog. The first 24 hours of continuous backup are included; retained backup beyond that window is a separate dimension. The phrase consumption pricing should not be paraphrased as paying only while an agent is actively working.
For budgeting, request a workload-shaped estimate rather than one blended capacity price. Include the primary-and-standby compute pair, any additional read or analytics nodes, logical data growth, retained history, and transfer. Then replay the intended workload through busy and quiet periods. The disclosed price-performance claim is useful for designing that test, but the retrieved billing page does not supply a universal dollar rate from which to calculate your entire migration saving.
Before pricing the trial, record logical data volume separately from allocated storage and query working set. The 32x nominal ceiling ratio cannot determine a migration budget because those quantities answer different questions. Ask the supplier to map the existing dataset into its billable units, including retained change history, and verify the resulting estimate against an actual invoice from the test. This is particularly important when a slide presents capacity and price-performance improvements together: neither figure supplies the missing conversion between your current storage accounting and the new bill.
There is also a less disruptive alternative for some buyers. Atlas Gen2 on GCP allows Standard IOPS to scale independently of storage capacity on eligible M30-plus dedicated clusters. That is not Infinite and should not be described as equivalent. It does mean that a team trying to solve an IOPS-overprovisioning problem should compare the upgrade already available on its platform before committing to a different database edition.
A familiar driver cannot carry the recovery plan
The strongest migration warning sits outside the query API. Infinite’s overview says the preview cannot live-migrate or restore backups between Atlas Core and Infinite, and a cluster’s edition cannot be changed after creation. MongoDB’s promise of no application-code changes concerns protocol and API compatibility. It does not supply a supported in-place conversion, a rollback path, or a recovery plan across the two editions.
The architecture documentation describes two electable compute nodes: a primary and a standby. The storage layer coordinates failover rather than a vote among compute nodes. Read preferences still work, but the document warns that reads sent away from the primary can concentrate on the single standby unless additional read-only or analytics nodes are configured. Keeping the connection code unchanged can therefore preserve syntax while changing the operational distribution of work.
Recovery deserves its own rehearsal. The restore guide excludes restores to another region, cloud provider, or organization during preview. Key management adds further conditions. A successful same-region restore is useful evidence, but cannot establish regional disaster recovery. If a production requirement depends on geographic recovery, leave that workload on a supported architecture until the required scenario is actually available and tested.
Encryption is not the missing feature. Atlas’s encryption overview says Infinite encrypts data on the compute node before it reaches shared storage. Customer-managed keys are supported through AWS KMS in preview. The same documentation ties restoration to key scope and availability. Treat those controls as part of the recovery design rather than assuming encryption and recoverability are interchangeable forms of protection.
Release control is another operating cost. Infinite uses the latest MongoDB 9.x release with automatic upgrades during preview; customers cannot choose a specific version. A team whose release process requires explicit certification before production upgrades should test whether that policy fits its requirements. The correct response is not to invent a workaround in the business case, but to keep incompatible workloads out of the trial’s deployment scope.
There are benefits available without conflating these decisions. MongoDB 9.0’s Queryable Encryption update adds prefix, suffix, and substring queries on encrypted fields. The update specifies bounded substring lengths and field sizes, so it is not arbitrary encrypted search. Evaluate that capability on its own terms rather than treating a server release, an encrypted-query feature, and a storage-edition preview as one indivisible purchase.
Our Sonnet cache analysis found that request compatibility and billing economics must be tested together. The database version is broader: a familiar driver can preserve the application contract without preserving the operating contract. Migration effort lives in that difference, even when no query needs rewriting.
Buy the preview as an experiment, not a destination
The bullish case is substantial. Separating storage durability and backup work from compute can make scaling and recovery less dependent on copying data among application-serving nodes. MongoDB reports customer examples alongside its internal benchmark, including a PicPay test at four times normal peak traffic for two hours. Those are vendor-published observations, not guarantees for another application’s write mix, working set, or service target. Use them to justify an experiment rather than to fill missing cells in a forecast.
The skeptic’s strongest case is not that the architecture cannot work. It is that the launch’s unified-platform story can conceal several independent readiness decisions. A retrieval-dependent application, a regulated recovery requirement, and a storage-heavy transactional service may reach different conclusions from the same documentation. A single enterprise commitment can simplify purchasing without making those conclusions converge.
Today’s Cribl analysis asks whether a cheaper model route preserves the accepted diagnostic result. Infinite asks whether a different storage route preserves the accepted operating result. In both cases, continuity at the API boundary is necessary but insufficient. Request records that show what changed underneath: billing dimensions, routing or topology, failure behavior, and the configuration that produced the measured outcome.
The Kumo brief likewise distinguishes a benchmark’s dataset ceiling from the context used for prediction. Capacity numbers are valuable when their definitions survive the purchase order. They become dangerous when a logical-storage limit becomes an assumed dataset multiplier, or when a maximum throughput specification becomes an expected production response time.
Build the trial around a workload that is eligible now. Preserve the incumbent’s service while measuring ingestion, steady-state requests, burst behavior, read distribution, failover, restoration, and the complete bill. Record any manual operational work as migration cost. Do not use production-only data or permissions merely to make a preview demonstration look representative; construct an authorized dataset and access boundary suitable for the trial.
Change the verdict when the evidence changes. Reproduced lower cost at the required latency would strengthen the case for a compatible transactional workload. Supported Vector Search, cross-edition migration, appropriate recovery options, and an uptime commitment would broaden the eligible audience. A larger storage ceiling alone would not settle any of those questions. Nor would a successful benchmark excuse a recovery rehearsal that failed.
- Database owners: test Infinite where storage growth and compute demand diverge, but keep workloads requiring unsupported search, regions, or recovery on a supported edition.
- Finance and platform teams: price provisioned compute, logical storage, backup retention, transfer, and parallel operation. Validate the entire bill before treating the 189% vendor benchmark as a saving.
- Agent leads: evaluate Agent Engine’s memory, runtime, and governance independently. Require an explicit compatibility map and accepted-task evidence before coupling its adoption to Infinite.
MongoDB has made the infrastructure choice more interesting, not universal. The operator advantage comes from selecting the new layer that solves a measured constraint while refusing the migrations that the preview cannot yet support.
Sources
- MongoDB — September 29 server and Infinite launch, readiness and benchmark claims
- MongoDB — Infinite capacity, supported features and preview exclusions
- MongoDB — AWS standard storage ceilings and Infinite throughput by tier
- MongoDB — Atlas Agent Engine modularity, pricing structure and preview
- MongoDB — Infinite compute autoscaling and storage-performance coupling
- MongoDB — Infinite compute, logical storage, backup and transfer billing
- MongoDB — GCP Gen2 independent Standard IOPS configuration
- MongoDB — Infinite topology, failover and read distribution
- MongoDB — Infinite restore support and geographic limitations
- MongoDB — encryption and customer-managed key boundaries
- MongoDB — Infinite version selection and automatic upgrades
- MongoDB — Queryable Encryption’s new supported query types