AWS said on 26 August it is working to bring NVIDIA’s Vera CPU to its cloud. NVIDIA’s own comparison table against Grace, published in the company’s technical blog on the Vera Rubin platform, is 88 custom Olympus cores, 176 threads of Spatial Multithreading, 2 MB of L2 per core, 164 MB of unified L3, up to 1.5 TB of LPDDR5X at up to 1.2 TB/s, and 1.8 TB/s of coherent NVLink-C2C to the Rubin GPUs. PCIe Gen6. CXL 3.1. Confidential computing, which Grace did not have, is on.

The underrated chip in the Rubin generation is not the GPU. It is the host that refused to stay a host.

That is silicon with a job. NVIDIA’s blog is blunt about the job: GPU performance alone does not keep a factory fed. Data, memory, and control have to move, or the expensive die sits idle. Vera is designed as a data engine rather than as a general-purpose apology sitting across a PCIe gap. AWS, on 26 August, gave that socket a landlord. The parts list is NVIDIA’s.

A dark liquid-cooled compute rack
A liquid-cooled compute rack in a dark hall. Not a photograph of a real event.

From Grace to Vera

NVIDIA Grace was 72 Arm Neoverse V2 cores. Threads matched cores: 72. L2 was 1 MB per core. Unified L3 was 114 MB. Memory was up to 480 GB of LPDDR5X at up to 512 GB/s. NVLink-C2C to the GPU was 900 GB/s. PCIe was Gen5. Confidential compute was listed as not applicable. SIMD was 4x 128-bit SVE2.

Vera is not a refresh sticker on that table. NVIDIA’s published comparison, in the same technical blog, puts 88 NVIDIA-designed Olympus cores on the die. Threads rise to 176 because each core runs two hardware threads under Spatial Multithreading, which is not the old time-slice. L2 doubles to 2 MB per core. Unified L3 rises to 164 MB. Memory bandwidth is listed at up to 1.2 TB/s, which NVIDIA calls 2.4x Grace. Capacity is up to 1.5 TB of LPDDR5X, which the same table calls 3x. NVLink-C2C doubles to 1.8 TB/s. SIMD is 6x 128-bit SVE2 with FP8. The I/O column moves to PCIe Gen6 and CXL 3.1. Confidential computing is a supported row.

Those are NVIDIA’s numbers. The figures are as NVIDIA printed them. Preliminary, the NVL72 product page says of the rack table. Subject to change. The shape of the argument is not going to change with a footnote. The CPU is no longer a polite x86 box on the other side of a slot. It is on the board, coherent, and fat enough in memory that the GPU can treat LPDDR and HBM as one address space.

Grace established the method: a high-bandwidth, energy-efficient CPU designed next to the GPU rather than borrowed from a server catalogue. Vera is that method with more cores, more memory, a wider coherent pipe, and a security row that Grace left blank. The cores are custom. Olympus, NVIDIA named them. Arm-compatible. The blog is specific: Arm v9.2, and major Linux distributions, frameworks, and orchestrators run unmodified. That is the sentence that keeps this from being a science project.

Arm spent 24 August at Hot Chips selling AGI, a finished 136-core server CPU of its own. That is on the record. Vera is a different product for a different socket. AGI is a host you can buy from Arm. Vera is a host NVIDIA built to sit 1.8 TB/s away from Rubin. Both are Arm. They are not the same chip.

Spatial Multithreading is the thermal paragraph

Olympus is a wide, deep core. NVIDIA’s blog lists improved branch prediction, prefetching, and load-store performance, aimed at control-heavy and data-movement work rather than at compiling a kernel for its own sake. Eighty-eight of them on one monolithic compute die. The thread story is the part that is new as a name.

Spatial Multithreading, as NVIDIA describes it, runs two hardware threads per core by physically partitioning resources instead of time-slicing them. The company says that lets the core make a run-time trade between performance and efficiency. Throughput and virtual CPU density go up. Isolation holds. Predictable latency is the point NVIDIA keeps returning to, because a multi-tenant rack is a place where one noisy neighbour used to steal the pipeline for a slice and give it back later. Partitioning is a different promise. Two threads, two sets of pipes, a switch you can make while the job is running.

176 threads from 88 cores is the arithmetic. Grace had 72 threads from 72 cores. The extra 104 threads are not free. They are the partition. The thermals are NVIDIA’s. NVIDIA did not publish a TDP for Vera as a standalone socket in the comparison table. The story it is selling is efficiency at the rack, not a tower wattage sticker. Fair enough, if you are buying racks. The partition is still the thermal paragraph: you can run the core as two quieter threads or as one louder one, at runtime, without pretending time-slicing is isolation.

Isolation matters in a multi-tenant rack. Predictable latency matters more. KV-cache offload, orchestration, agentic traffic, data staging, scheduling. Those are the jobs NVIDIA lists for Vera as a host and as a standalone platform. The GPU still does the matrix. The CPU keeps the GPU from waiting.

A dark CPU package on a tray
A server CPU package on a black tray. Not a photograph of a real event.

One die, one fabric, swap the memory

The second-generation Scalable Coherency Fabric is the quiet row. NVIDIA says SCF connects all 88 Olympus cores to a shared L3 and to the memory subsystem on a single monolithic compute die. No chiplet boundary in the middle of that conversation. The company claims the fabric sustains more than 90 percent of peak memory bandwidth under load. Linear scaling of orchestration as core count rises is the sentence that follows. The fabric is a claim until the rack is full. The design choice is still the design choice: keep the 88 cores on one die so a request does not take the scenic route through a stitch.

Memory is LPDDR5X, up to 1.5 TB, up to 1.2 TB/s, at low power relative to a forest of DDR5 DIMMs. The serviceability story is SOCAMM: Small Outline Compression Attached Memory Modules. NVIDIA’s blog says the modules improve serviceability and fault isolation. Swap memory without treating the whole CPU as a brick. Grace did not have that sentence. In a rack that is supposed to run without a maintenance window, a memory module you can pull is worth more than a brochure about uptime. Vera CPUs, the same blog says later, complement GPU-level resiliency with in-system core validation, shorter diagnostics, and those SOCAMM modules.

Second-generation NVLink-C2C is 1.8 TB/s of coherent bandwidth between Vera and Rubin. Unified address space. Applications can treat LPDDR5X and HBM4 as a single coherent pool. NVIDIA names the uses: KV-cache offload, multi-model execution, less copying. The number to hold onto is the doubling. Grace was 900 GB/s. Vera is 1.8 TB/s. The GPU is no longer shouting across a PCIe gap and hoping the host notices.

PCIe Gen6 still exists. CXL 3.1 still exists. Those are the rows for the rest of the world: NICs, storage, whatever still speaks the old bus. The interesting pipe is C2C. The interesting pool is 1.5 TB of LPDDR next to however much HBM4 the Rubin dies are carrying. NVIDIA’s GPU table, same blog, puts up to 288 GB of HBM4 on each Rubin with aggregate bandwidth up to 22 TB/s. The NVL72 product page’s rack table lists 20.7 TB of HBM4 across 72 GPUs. 288 GB times 72 is 20.736 TB. The arithmetic holds.

Rubin, because the CPU has a job

A CPU story that never looks at the GPU is a brochure. Rubin is why Vera is 88 cores instead of 8.

The Rubin GPU, on NVIDIA’s blog, integrates 224 streaming multiprocessors. Fifth-generation Tensor Cores. NVFP4 and FP8. Transistor count for the full chip is 336 billion, against Blackwell’s 208 billion. Two compute dies, as Blackwell had. NVIDIA’s compute table lists 50 PFLOPS of NVFP4 inference on Rubin against 10 on Blackwell, and 35 PFLOPS of NVFP4 training against 10. Softmax acceleration in the SFU column doubles the FP32 and FP16 ops-per-clock-per-SM figures NVIDIA printed for Blackwell. Those are NVIDIA’s peak columns. The blog is careful, in other rows, to talk about sustained throughput. The distinction stays.

HBM4 doubles interface width against HBM3e, NVIDIA says, and nearly triples memory bandwidth against Blackwell. Up to 288 GB per GPU. Up to 22 TB/s. Decode and front-end efficiency are the sentences NVIDIA attaches to long-context inference. Scientific computing is still in the table: FP32 vector 130 TFLOPS, FP32 matrix 400 TFLOPS using Tensor Core emulation, FP64 vector 33 TFLOPS, FP64 matrix 200 TFLOPS by the same emulation note. Hopper had dedicated FP64 matrix hardware. Blackwell and Rubin, NVIDIA writes, do high FP64 matrix throughput with multiple passes over lower-precision tensor cores. Vector FP64 remains for the codes that are not matrix kernels and are bound by registers, caches, and HBM. That is a real distinction. It is also a GPU paragraph. Vera’s job is to keep those SMs from starving while the precision argument goes on.

NVLink 6 is 3.6 TB/s of bidirectional GPU-to-GPU bandwidth per GPU, double Blackwell’s 1,800 GB/s. NVLink-C2C CPU-to-GPU is 1.8 TB/s, double Blackwell’s 900. PCIe stays 256 GB/s bidirectional Gen6 on both. Inside an NVL72, NVIDIA describes a full all-to-all across 72 GPUs. MoE routing, collectives, synchronization-heavy inference. The switch tray is the other chip. Hold that.

The third-generation Transformer Engine, NVIDIA says, adds hardware-accelerated adaptive compression for NVFP4 and is compatible with Blackwell, so previously optimized code moves. 50 petaflops of NVFP4 for inference is the figure the CES press release on 5 January repeated. Same number as the blog’s GPU table. Same number as the NVL72 product page’s per-GPU column.

Superchip, then tray, then rack

NVIDIA’s building block is not a dual-socket motherboard from 2019. It is the Vera Rubin superchip: two Rubin GPUs and one Vera CPU, tied with memory-coherent NVLink-C2C, on one host processing motherboard. The CPU is in the execution path. Data movement, scheduling, synchronization. Not an external host.

The NVL72 product page prints the superchip column next to the rack and the single GPU. NVFP4 inference: 100 PFLOPS on the superchip, 50 on the GPU, 3,600 on the rack. NVFP4 training: 70, 35, 2,520. FP8/FP6 training: 35, 17.5, 1,260. CPU cores: 88 on the superchip, 3,168 on the rack. CPU memory: 1.5 TB LPDDR5X on the superchip, 54 TB on the rack. 36 times 1.5 is 54. 36 times 88 is 3,168. 36 Vera CPUs and 72 Rubin GPUs is the NVL72 configuration NVIDIA has been repeating since January. Two GPUs per CPU. That is the superchip, times 36.

NVLink bandwidth on that table: 7.2 TB/s on the superchip, 3.6 per GPU, 260 TB/s of NVLink 6 switch bandwidth on the rack. CES said 260 TB/s for the rack and called it more bandwidth than the entire internet. NVIDIA’s sentence. NVLink-C2C: 1.8 TB/s on the superchip, 65 TB/s on the rack. Scale-out networking: 0.8 TB/s on the superchip, 0.4 TB/s per GPU, 28.8 TB/s on the rack. 1.6 Tb/s per GPU from ConnectX-9 is 0.2 TB/s one way; bidirectional 0.4; times 72 is 28.8. The arithmetic holds again.

Total NVIDIA plus HBM4 chips: 12 on the GPU column, 30 on the superchip, 1,296 on the rack. Preliminary, the footnote says. All values up to, and subject to change.

The compute tray takes two superchips. Power, cooling, networking, management, modular, cable-free. NVIDIA says a redesigned liquid manifold and universal quick-disconnects support higher flow than the prior generation. Independent front and rear bays. The tray goes offline for maintenance. Assembly that took more than 1.5 hours on Blackwell is listed at about five minutes on Vera Rubin. Up to 18x faster, the blog and the CES release both say. There is no stopwatch reading from that aisle here. The claim is the claim. Cable-free is the part that will either survive a Tuesday or not.

ConnectX-9 SuperNICs in the tray: 1.6 Tb/s per GPU. BlueField-4 DPUs offload networking, storage, and security. The GPUs stay on the matrix. The CPU stays on the data. The DPU stays on the plumbing.

A dark GPU and CPU board
CPU and GPU silicon under a cooling plate. Not a photograph of a real event.

The other chips in the box

The Vera Rubin platform, as NVIDIA described it on 5 January at CES, was six new chips. The technical blog still opens with six, then notes an update on 16 March 2026: a seventh, NVIDIA Groq 3 LPX, as a low-latency inference accelerator for the same platform. The NVL72 product page now leads with seven. The six that Vera actually talks to are named, then the seventh as NVIDIA named it.

NVLink 6 switch: sixth-generation scale-up fabric, 3.6 TB/s GPU-to-GPU per GPU. Switch trays form an all-to-all across the rack. NVIDIA SHARP in-network compute sits in the fabric. The blog puts 14.4 TFLOPS of FP8 in-network compute on each NVLink 6 switch tray. All-reduce, reduce-scatter, all-gather can run in the switch. NVIDIA says that can cut all-reduce traffic by up to 50 percent and improve tensor-parallel execution time by up to 20 percent, depending on model, parallelism, participant count, and NCCL. Hot-swappable trays. Partial racks. Dynamic reroute when a switch goes offline. In-service software updates. Telemetry. Operable, not just fast. That is the RAS paragraph for the fabric.

ConnectX-9: the endpoint. Quad SuperNIC boards in the compute tray. 1.6 Tb/s per Rubin GPU. Programmable congestion control, traffic shaping, packet scheduling at the endpoint, in concert with Spectrum-6. IPsec and PSP for data in transit. Encryption for data at rest. Secure boot, firmware authentication, device attestation. NVIDIA’s security list, not a lab result.

BlueField-4: a dual-die package. 64-core Grace CPU for infrastructure offload, plus ConnectX-9, up to 800 Gb/s Ethernet or InfiniBand. Memory 128 GB at 250 GB/s, against BlueField-3’s 32 GB at 75 GB/s. Hosts: 128K against 32K. NVMe-over-fabrics: 20 million IOPS at 4K against 10 million. DOCA is the software layer. ASTRA is the trust architecture NVIDIA introduced with BlueField-4 on this platform: a trust domain in the compute tray, control of provisioning and isolation off the host. Inference Context Memory Storage, ICMS, is the storage sentence. BlueField-4 runs a KV-cache tier NVIDIA calls G3.5, Ethernet-attached flash between HBM/DRAM/local SSD and durable shared storage. NVIDIA claims up to 5x tokens per second and up to 5x power efficiency against traditional storage approaches for that KV path. Dynamo and NIXL orchestrate the blocks. Spectrum-X carries the RDMA. This is NVIDIA describing its own stack. Put it in the file. Do not confuse it with a crate that has already been billed.

Spectrum-6: 102.4 Tb/s per switch chip, 128 times 800 Gb/s, 200G PAM4 SerDes. Spectrum-4 on Grace Blackwell was 51.2 Tb/s, 64 times 800 Gb/s, 100G PAM4. Co-packaged optics. NVIDIA’s photonics claims on the product page: 5x better network power efficiency, 10x higher network resiliency, up to 5x more uptime over pluggable transceivers. The blog puts optical loss at about 4 dB against about 22 dB, and talks about 64x better signal integrity. MMC-12 cabling. Spectrum-XGS for distance-aware congestion control across sites. Standards-based Ethernet, NVIDIA still says, with the behaviour tuned for bursty all-to-all rather than for office mail.

Groq 3 LPX, from the NVL72 page: 256 LPUs per rack, 128 GB SRAM, 40 PB/s memory bandwidth, 640 TB/s scale-up bandwidth per rack. Co-designed with NVL72. NVIDIA’s figures for that pairing include up to 35x inference performance per watt and up to 10x more revenue opportunity for trillion-parameter models relative to Blackwell, on NVIDIA’s own tier table. Projected, subject to change, the footnotes say. Vera does not become a Groq chip because a seventh die joined the poster. Vera is still 88 Olympus cores. The LPX rack is the low-latency neighbour.

There is also HGX Rubin NVL8, from the 5 January CES release: eight Rubin GPUs through NVLink, for x86-based platforms that still want a conventional host. And Vera Rubin NVL4, on the product page: four Rubin GPUs on a second-generation NVLink bridge, two Vera CPUs over NVLink-C2C, liquid-cooled MGX servers, with NVIDIA’s claimed gains against Grace Hopper of up to 4x for scientific simulation, 6x for AI-for-science training, and 8x for AI-for-science inference. Smaller boxes. Same coherent idea. The story is the 88-core socket. The NVL4 and NVL8 SKUs are how that socket shows up when you are not buying 36 of them at a time.

5 January, a mountain, and the second half of the year

NVIDIA launched the Rubin platform on 5 January 2026 at CES in Las Vegas. Six chips, one supercomputer, the press release said. Named for Vera Florence Cooper Rubin, the astronomer. The CPU is named Vera. The cores are named Olympus. The cores are the story.

The CES text put Rubin in full production and said Rubin-based products would be available from partners in the second half of 2026. First cloud names, in that release: AWS, Google Cloud, Microsoft, OCI, and NVIDIA Cloud Partners CoreWeave, Lambda, Nebius, and Nscale. Microsoft’s Fairwater superfactories were named as a Vera Rubin NVL72 destination, at a scale NVIDIA described as hundreds of thousands of Vera Rubin Superchips. CoreWeave said it would integrate Rubin-based systems in the second half of 2026. Cisco, Dell, HPE, Lenovo, and Supermicro were listed as expected server vendors. Red Hat announced an expanded stack on the same day: Enterprise Linux, OpenShift, Red Hat AI.

The CES numbers for the platform against Blackwell were the ones the later blog and the product page still carry: up to 10x reduction in inference token cost, and 4x fewer GPUs to train MoE models. NVLink 6 at 3.6 TB/s per GPU and 260 TB/s per NVL72 rack. Vera at 88 Olympus cores, Armv9.2, NVLink-C2C. Rubin GPU at 50 petaflops NVFP4 inference. Confidential computing across CPU, GPU, and NVLink, which CES called a first at rack scale for NVIDIA. Second-generation RAS Engine. Cable-free trays, up to 18x faster assembly than Blackwell.

The technical blog, later, walked those claims with tables. Training a 10-trillion-parameter MoE on 100 trillion tokens in a one-month window, NVIDIA’s figure, uses about one-quarter the GPUs on Vera Rubin NVL72 versus Grace Blackwell NVL72. Inference on Kimi-K2-Thinking at 32K/8K input/output, NVIDIA’s figure, is up to 10x more tokens per megawatt than GB200 NVL72 at comparable interactivity, and up to 10x lower cost per million tokens. The NVL72 product page repeats both, with the usual “projected, subject to change.” There is no 10-trillion-parameter training run behind this report. It has NVIDIA’s comparison, dated, named, and consistent across the CES release, the technical blog, and the product page.

The NVL72 product page, as of 30 August, says Vera Rubin is ramping into full production, with Taiwan’s top server makers and the supply chain manufacturing at scale and shipping systems. That is NVIDIA describing its own ramp. Japan is a named customer on the same page: NVIDIA working with Noetra Corp. on a Vera Rubin installation of 13,750 Vera CPUs and 27,500 Rubin GPUs. 13,750 times two is 27,500. Superchip arithmetic again. Working with, not a ribbon-cutting attended here.

NVIDIA DGX SuperPOD is the deployment unit in the blog: eight DGX Vera Rubin NVL72 systems, Spectrum-X scale-out, Mission Control, certified storage. MGX third-generation rack, same footprint as the prior generation, more than 80 partners. 3.6 exaFLOPS of AI inference compute per NVL72 rack is the density sentence on the blog. 3,600 PFLOPS NVFP4 inference on the product-page table is 3.6 exaFLOPS. Same number, two pages.

AWS, 26 August

The news this night is not the January keynote. It is the landlord.

On 26 August 2026, Amazon Web Services and NVIDIA put out a joint release from Seattle and Santa Clara. The headline was 2 million additional GPUs and next-generation infrastructure for agentic and physical AI. Under that headline, a specific CPU sentence: AWS and NVIDIA are working to bring Vera CPU-based infrastructure to AWS, as an option for workloads that want high-performance CPU next to the accelerators. Purpose-built for the next generation of AI, the release said, and complementary to AWS’s custom silicon. Matt Garman, CEO of AWS, talked about choice and about NVIDIA on AWS. Jensen Huang, founder and CEO of NVIDIA, talked about 16 years of scaling NVIDIA in the cloud together, and about expanding the partnership across GPUs, CPUs, networking, open models, and software.

The GPU count in that release is separate from the Vera sentence and should stay separate. At GTC 2026, AWS had already said it would add more than 1 million NVIDIA GPUs starting in 2026. On 26 August AWS said demand had exceeded that, and that it plans to deploy an additional 2 million NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra GPUs in 2027 and 2028 across AWS global infrastructure, including AI factories. That is a plan, dated, with SKU names. It is not a Vera crate count. Vera, in the same release, is “working to bring.” There is no socket number AWS did not print.

The rest of the 26 August list is the rest of the stack, not a reason to bury the CPU. NVLink Fusion, already announced for next-generation Trainium at re:Invent 2025, expands to NVIDIA’s custom high-bandwidth memory, NVHBM, with Annapurna Labs and memory suppliers. Trainium and GPUs in a common rack-scale architecture is the aim as written. AI factories for the U.S. government, including plans for 100,000 GPUs on secure AWS infrastructure for federal and national-security workloads at Impact Level 6 and above. All NVIDIA GPU and Trainium EC2 instances, including NVLink Fusion, built on the Nitro System and scaled out through Elastic Fabric Adapter. Nemotron open models on Amazon Bedrock as managed serverless, and on SageMaker for self-deploy. GPU-accelerated EMR with EC2 G7 and cuDF, AWS saying up to 3.7x faster processing and 30 percent better price-performance than CPU configurations. GPU-accelerated vector indexing on OpenSearch, AWS saying up to 9x faster indexing at a quarter of the cost. Amazon Robotics on Jetson, Omniverse, and Isaac. G7 instances with RTX PRO 4500 Blackwell Server Edition, AWS saying it is the first major cloud to offer that SKU, at 4.6x AI inference and 2.1x graphics against G6.

Spectrum networking is in the AWS release too. So is the 16-year partnership. So are the forward-looking-statement footnotes, which are long, and which are not a ship date.

Among the first clouds to deploy Vera Rubin-based instances in 2026, the January CES release had already named AWS. The August release is the CPU-specific follow-through: not only Rubin GPUs on AWS, but Vera as a CPU option beside them. That is the socket finding a landlord. It is still not a crate count.

Software that does not have to be rewritten

Vera supports Arm v9.2. NVIDIA says major Linux distributions, AI frameworks, and orchestrators run unmodified. CUDA backward compatibility across hardware generations is the GPU-side version of the same idea. NeMo, Megatron Core, TensorRT-LLM, vLLM, SGLang, Dynamo, NCCL, NIXL, cuDNN, CUTLASS, FlashInfer. NVIDIA AI Enterprise as the supported stack. Mission Control for the factory: Run:ai workload management, power profiles, autonomous recovery, dashboards, health checks. This is a lot of product names. The hardware point is smaller. If the 88 cores speak Arm v9.2 and the distributions boot, the rest of the stack is a software desk’s problem. NVIDIA is not asking anyone to recompile the world for Olympus. That is the compatibility sentence that matters.

Confidential computing is the other software-shaped hardware row. Grace did not have it. Vera does. The blog describes Vera Rubin NVL72 as a rack-scale trusted execution environment across CPUs, GPUs, and interconnects: TDISP, PCIe IDE, NVLink-C2C encryption, NVLink encryption. Remote attestation through NVIDIA’s NRAS, on-demand or cached, including air-gapped. Near-native performance is NVIDIA’s claim. Hopper started GPU confidential computing. Blackwell removed some of the performance tax. Rubin, NVIDIA says, unifies CPU and GPU into one trust domain across the rack. Third-generation, in their numbering. The generation count is NVIDIA’s. The useful fact is that the CPU column finally has a yes.

Heat, water, and what lasts

AI factories, NVIDIA writes, can lose about 30 percent of incoming power to conversion, distribution, and cooling before a GPU does any work. Parasitic energy. Vera Rubin NVL72, like Blackwell, uses warm-water single-phase direct liquid cooling at a 45-degree Celsius supply. Ambient-air cooling of that water is the point of staying at 45 rather than dropping to 35. The blog says Vera Rubin nearly doubles thermal performance in the same rack footprint without a new cooling scheme. Rack-level power smoothing, with about 6x more local energy buffering than Blackwell Ultra, is the answer NVIDIA gives to megawatt-scale ramps during all-to-all. Site-level batteries for grid events. Domain Power Service, Workload Power Profile Solution, SMI, DCGM. Omniverse DSX for grid-aware control: DSX Flex to translate utility signals, DSX Boost to enforce ramp rates. NVIDIA says that combination can put up to 30 percent more GPU capacity in the same power envelope. A Vera Rubin NVL72 research centre in Manassas, Virginia, is named as the place NVIDIA is validating 100 MW up to gigawatt-scale reference designs.

Those are facilities numbers. This is a hardware story. The thermal paragraph that belongs here is simpler. 45-degree water. Modular trays. SOCAMM memory you can swap. A fabric NVIDIA says keeps more than 90 percent of peak memory bandwidth. A CPU that can partition a core instead of pretending a time-slice is isolation. Assembly in minutes rather than in an hour and a half, if the 18x claim holds on a real floor. Second-generation RAS on the GPU: health checks in idle windows, in-field SRAM repair, zero-downtime self-test. NVLink Intelligent Resiliency so a rack can run while a tray is pulled. Cable-free, hose-free, fanless compute tray against the cabled Blackwell Ultra tray, in NVIDIA’s own comparison animation.

What lasts, on this beat, is the boring bit. A memory module that comes out. A tray that comes out. A core that can be validated in place. A coherent pipe that does not drop to PCIe when the GPU asks for a page. Eighty-eight cores on one die so the fabric does not negotiate a chiplet. Arm v9.2 so the disk image still boots. Confidential computing so the row in the table is no longer NA.

The NVL72 table is preliminary. The AWS Vera sentence is “working to bring.” The 2 million GPUs are 2027 and 2028, Blackwell Ultra and Rubin and Rubin Ultra mixed. The 10x inference and 4x training figures are NVIDIA against NVIDIA, on named model setups, projected. Japan’s 13,750 CPUs are a working-with. H2 2026 availability was the January line; August’s product page says shipping. Those are not flattened into a single crate. The chip in the table is still the chip in the table.

This desk

Fujitsu put Monaka’s cache on its own 5 nm die. Arm sold a finished 136-core AGI at Hot Chips on 24 August. XMG put 240 Hz OLED on a 40.6 cm aluminium box. Lenovo shipped SteamOS on a handheld. Those were this week’s shorter filings. Vera is the longer one because the socket changed shape.

NVIDIA built a custom Arm core, put 88 of them on one die, doubled the coherent pipe to the GPU, tripled the LPDDR, and asked the host to stop being a host. AWS, on 26 August, said it will put that CPU in the cloud next to the accelerators. The cores are named Olympus. The astronomer is named Vera Rubin. The rack is named NVL72. The number that will still be true when the posters change is 88.

Eighty-eight cores, a 1.8-terabyte-a-second pipe, and a memory pool you can treat as one address space. That is a good chip. The next facts to watch are the trays coming out of Taiwan, the first Vera instance on an AWS price list, and whether Spatial Multithreading holds when the rack is loud. Until then, the table is the table. Print it.