Compute & Market Power
Brazil's G4 Brings 4x Memory, Not Blanket Residency
Google brings G4 GPUs to São Paulo, with four times L4 memory per GPU. Gemini's October 15 processing promise remains product- and model-specific.
Brazilian infrastructure teams can now qualify Google Cloud’s G4 virtual machines in São Paulo, announced September 24. A full G4 GPU offers 4x the memory of an L4 GPU in Google’s G2 series, but the same announcement’s October 15 Gemini processing commitment does not make every Google AI product locally resident today.
A larger memory budget is not a faster-token promise
The hardware comparison is unusually clean if its scope stays narrow. NVIDIA specifies 96 GB of GDDR7 memory for the RTX PRO 6000 Blackwell Server Edition. Google’s accelerator documentation lists 24 GB per L4 GPU in the G2 series. Divide 96 by 24: 4x per physical GPU. These are published capacities, not independently measured application results.
A full G4 GPU carries four times the memory of an L4
GPU memory per full physical GPU, GB. Capacity, not throughput or price.
That gap is useful for teams whose workload does not fit comfortably on one L4. More memory can change the set of candidate deployments, but it cannot establish which one is economical. The comparison says nothing about tokens per second, concurrent users, effective model quality, or price per accepted request. Those quantities depend on the model and serving configuration the team actually runs.
Even the memory figure needs the correct purchasing unit. Google’s current machine table includes fractional G4 configurations as well as full GPUs. Its full single-GPU g4-standard-48 shape pairs the GPU with 48 vCPUs and 180 GB of instance memory. The 96 GB figure refers to GPU memory, not that system RAM. A procurement request for a G4 instance without specifying its shape can therefore fail to buy the capacity described by the headline.
The regional news is availability, not a new accelerator invention. Google says these machines are now available in the São Paulo cloud region. That gives locally situated workloads a candidate worth testing. It does not guarantee a particular account’s quota, reserved supply, or application latency. Ask the account team to prove the intended shape can be provisioned in the required zone before attaching a production date to the announcement.
Migration has storage implications too. Google’s G4 documentation says Persistent Disk is not supported and identifies Hyperdisk Balanced as the supported boot-disk type. An existing GPU image or deployment plan may therefore need adaptation beyond changing an instance name. Include the disk configuration, data movement, operational recovery, and tooling changes in the evaluation rather than treating the GPU memory ratio as a complete migration specification.
There is no defensible São Paulo hourly price in the retrieved announcement. A rate for another region would not close that gap. Request a regional quote for the full configuration, including storage and the chosen consumption option. The useful economic denominator is accepted workload at the required latency and reliability, not memory alone. A larger card can be rational even at a higher hourly price, but only if the workload earns that premium.
Residency comes with a product name and a date
The managed-service promise is separate. Google says that starting October 15, Gemini Enterprise on the web will support local in-country processing for Gemini 3.5 Flash in Brazil, alongside existing support for customer-data storage for agent workloads. That wording identifies a product surface, a model, and a future start date. It is not an assurance about every Gemini model, every API, or every feature used by an enterprise agent.
The documentation currently reinforces the need to verify scope. The Gemini Enterprise Standard and Plus residency page retrieved for this edition lists Canada, India, Japan, Singapore, and the United Kingdom as in-country locations; Brazil is not yet in that table. The launch announcement may precede a documentation update. Neither source should be silently rewritten to make them agree. Obtain the applicable configuration and commitment before declaring a Brazil deployment compliant with an internal residency requirement.
The same documentation distinguishes data stored at rest from machine-learning processing and describes feature-specific limitations. That is the right way to review the announced service. A locally stored document can still be used by a feature with different processing terms. Ask which model receives it, which endpoint is selected, and whether grounding, execution, or connected services carry separate constraints. Geographic proximity is not a substitute for a data-flow review.
Our Alibaba analysis separated an expanding regional footprint from service-specific delivery. Google’s announcement is more immediate for G4, but its managed-AI commitments still have their own clock. Self-hosting a workload on regional GPUs and using a managed Gemini application are different deployment decisions. Do not let a shared keynote turn them into one procurement approval.
The switching cohort is therefore split. Infrastructure teams constrained by GPU memory should test the full-GPU G4 shape against their current serving stack. Buyers waiting for local Gemini Enterprise processing should prepare a product-specific acceptance review for the announced start date. Neither group should wait for the other to finish, and neither should borrow the other’s evidence as proof of its own readiness.
Today’s Ando lead makes the same distinction between a product’s attractive premise and its operational dependencies. Here the verdict changes when quota, configuration, local workload measurements, and written residency scope line up. Expand if the measured application economics improve and the intended data flows are covered. Delay if a memory-rich machine leaves storage migration, regional pricing, or processing commitments unresolved. Four times the memory is a useful specification; it is not four times the permission to proceed.