Nvidia stood up at Hot Chips on 24 August and said NVIDIA Groq 3 LPX is in full production. The newsroom post is dated that Monday. The chip is the LP30. The rack holds 256 of them. Nebius is the first AI cloud on the list. The $20 billion Groq deal from December now has a crate.
Artificial Analysis ran Gemma 4 31B, an open agentic model, at a 100,000-token context and measured 3,431 output tokens per second. Nvidia’s own newsroom rounded that to 3,400 and called it the fastest figure ever recorded for the model. Tom’s Hardware, filing Luke James on 26 August, kept 3,431 and put the next-fastest public endpoint at 870 tokens per second, roughly four times slower. Those are the two numbers on the page. The rack is a decode box. It is not a training hall.


Igor Arsovski presented the architecture. He was Groq’s chief architect. He is now Nvidia’s vice president of hardware. Tom’s Hardware had him calling the morning “a pinch me moment for the Groq team that’s now integrated into the Nvidia group.” The LP30 carries roughly 500 megabytes of on-die SRAM and no HBM. A full LPX rack of 256 chips therefore holds 128 gigabytes of memory and, on Nvidia’s slide, 40 petabytes a second of aggregate bandwidth against 315 petaFLOPS of FP8 compute, with 350 nanoseconds of chip-to-chip latency. The rack is Vera Rubin-compatible, MGX, liquid-cooled, and Nvidia says it scales past 1,000 LPUs. SRAM instead of HBM is the whole trick. Weights sit on the die. The decoder does not wait on a stack.
CNBC, the same Monday, had Nvidia senior director Dion Harris telling reporters the Groq rack will sit beside Vera central processors and Rubin graphics processors at Nebius and will be online later this year. Harris’s useful sentence is the one that keeps the story honest: “This isn’t about replacing GPUs. It’s about using the right price, right processor for the right part of the workload.” Decode is the part. Prefill still likes a GPU. Groq chips are manufactured by Samsung. TSMC still makes Nvidia’s GPUs. CNBC printed both foundries. They are not swapped.
Nvidia’s newsroom put Jensen Huang on the Monday post. “Inference is the growth engine of AI,” he said. “NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency. Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.” The quote is a founder on a production day. The industrial fact under it is 256 LP30s in a liquid-cooled rack, in production, 24 August.
Nebius is first. Danila Shtan, the company’s chief technology officer, said generation is the phase of inference that determines how responsive a system actually is, and that Groq 3 LPX is built to accelerate that phase. Nebius Token Factory is the production inference platform. Developers keep the same API. “No migration to a new stack,” Shtan said. Following Nebius, Nvidia’s release said purpose-built inference cloud Groq plans to be among the platform’s earliest adopters. Groq the cloud and Groq the silicon are no longer the same company. The December deal was a non-exclusive IP licence plus the hiring of Jonathan Ross, president Sunny Madra, and most of Groq’s engineers. Tom’s Hardware noted that structure avoided a formal merger review. This is a production story, not a competition story. The licence is how the LP30 got an Nvidia name. Full production is how it left the slide.
Artificial Analysis’s 3,431 is a third-party number with a method. Tom’s Hardware said the comparison ran on a private, pre-release Gemma 4 31B endpoint served through Google Cloud, median of 50 sequential client requests at a concurrency of one, while the public providers in the same chart ran shared production serverless endpoints. One request at a time produces the highest per-user token rate the hardware can post. It is not a multi-tenant production mix. Nvidia’s on-stage demo showed 10,996 tokens per second on the same 31B model. Arsovski flagged that figure on stage as “self-reported” before telling the room the aim was “third-party verified independent benchmarks that you guys can trust.” 3,431 is the independent cell and 10,996 is the self-reported demo, labelled. Gemma 4 31B is a dense model small enough to sit inside a single LPX rack. A trillion-parameter mixture-of-experts is a different memory problem. Tom’s Hardware said that picture went unaddressed on stage. The missing picture stays missing.
SRAM is the constraint as well as the gift. A Rubin GPU carries 288 gigabytes of HBM4, roughly 576 times the memory of a single LP30. A 31-billion-parameter model at FP8 needs on the order of 62 LPUs to hold its weights. A large mixture-of-experts runs into four figures of chips across several racks. Capacity is the cost of keeping the working set on die. Nvidia is describing the LPU as a decode co-processor, not as a general-purpose replacement for its GPUs. That is Harris’s sentence made into a floorplan.
The pipeline is deterministic. The design drops caches, branch prediction, and out-of-order execution. The compiler schedules at clock-cycle granularity. The architecture comes from the Tensor Streaming Processor Groq described in a 2020 ISCA paper titled Think Fast, the same title Arsovski and Raghavan reused at Hot Chips. Ross founded Groq after Google’s TPU work. Determinism lets the compiler predict power draw cycle by cycle. Nvidia uses that to pre-order current from the rack’s regulators ahead of demand, cutting voltage droop by more than 60 percent and overshoot by more than 70 percent against an uncompensated load. The same per-block scheduling equalizes heat instead of throttling to the hottest tile. Arsovski put that at roughly 10 to 11 percent additional performance under a fixed thermal limit. Those are Nvidia’s thermal sentences, said on a Hot Chips floor, about a rack the company says is already being built.

Across racks, Nvidia synchronizes chips to a single virtual clock in what it calls a plesiosynchronous network. Each chip is both processor and router. The fabric does not need adaptive routing or congestion sensing. Clock drift is compensated at the chip-to-chip links. Asked about a chip that fails mid-workload, Arsovski said users “would experience the exact same as any other hardware in the industry” and would “just checkpoint it or reconfigure the hardware.” That is a reliability answer, not a reliability study. The next number is a mean-time-between-failure when Nvidia prints one.
The split with Vera Rubin NVL72 is the product. Rubin GPUs handle the compute-heavy prefill and build the KV cache. LPUs generate output tokens. Nvidia showed three ways to divide the work: disaggregated prefill and decode; attention-FFN disaggregation, which keeps attention and its cache on GPU HBM while the LPU runs the feed-forward layers; and external-draft speculative decoding, where a small model on the LPU proposes tokens that the GPU verifies in parallel. An FPGA bridges the synchronous LPU domain and the asynchronous world of host I/O and GPU hand-offs. Dynamo, plus an LPU extension to CUDA, orchestrates the split. Nvidia put the gains from those modes at roughly three-to-five times over Rubin alone on a two-trillion-parameter workload with a 400,000-token cached context. Those are Nvidia-measured. They are not promoted into Artificial Analysis cells. The independent cell remains 3,431 tokens per second on Gemma 4 31B.
Cerebras used the same Hot Chips session to present CS4, a wafer-scale decode pitch of its own, and it already has a July agreement to pair AMD Helios GPUs for prefill with its engines for decode. That is the other company in the room making the same argument: decode is a different job from prefill. Nvidia’s answer is in-house. Ian Buck said at GTC 2026 that Nvidia pulled Rubin CPX, its own GDDR7 long-context accelerator, to focus on shipping the LPU this year. The LPX rack is that shipment. 24 August is the production sentence.
Huang, at the Vera Rubin and Groq 3 LPX unveiling in March, projected $1 trillion in cumulative sales between current-generation Blackwell chips and the new Vera Rubin systems through 2027. He said he would allocate a quarter of data-centre space intended for coding applications to Groq chips, with the rest of that hall 100 percent Vera Rubin. CNBC reprinted those March lines next to Monday’s production news. A quarter of a coding hall is a founder’s allocation, not a 24 August packing list. What 24 August adds is the crate: 256 LP30s, Samsung silicon, 500 megabytes of SRAM each, 128 gigabytes in the rack, 40 petabytes a second, 315 petaFLOPS of FP8, 350 nanoseconds chip-to-chip, Nebius first, later this year on Harris’s calendar.
There is no customer invoice with a dollar on it. A street price for an LPX rack. A multi-tenant production mix that matches Artificial Analysis’s concurrency-of-one method. A claim that SRAM decode retires the GPU. Harris already declined that claim. OpenAI’s Jalapeño, 700 watts, sat in the same Hot Chips week as a different inference story. This page is the Nvidia crate.
The facts: 24 August, Hot Chips, Groq 3 LPX in full production, 256 LP30s, 500 megabytes of SRAM, no HBM, 3,431 tokens per second on Gemma 4 31B from Artificial Analysis, 10,996 self-reported on stage, four times the 870-token public endpoint in Tom’s Hardware’s chart, Nebius Token Factory first, Samsung on the Groq wafer, TSMC still on the GPU, $20 billion in December, a decode co-processor bolted to Vera Rubin NVL72, later this year for the first cloud. Named rack. Named date. Named third-party bench. The LPU rack is in production. That is a factory sentence, and it is this week’s.

The paper
Comments
No notes on this story yet.
Sign in to comment