BEP Research

BEP Research

NVIDIA Vera: When CPU Latency Becomes GPU Economics

NVIDIA named AMD’s flagship as the yardstick — and quantified the chiplet tax.

Ben Pouladian's avatar
Ben Pouladian
Jul 21, 2026
∙ Paid

“Chiplet architectures are really becoming the legacy CPU architecture... chiplets levy a heavy tax on memory bandwidth and data movement.”

That was the closing statement at the Vera analyst briefing last Thursday, July 16, where I attended live and asked a question. “Chiplet tax” was NVIDIA’s phrase on the closing slide, aimed at the design philosophy under every EPYC server CPU AMD has shipped since 2017 and a growing share of Intel’s Xeon line — filed under legacy, on the record. This morning NVIDIA published the whitepaper that prices the claim: the first NVIDIA CPU document I can find that prints a current competitor’s flagship part number beside NVIDIA’s own score and publishes the full comparison configuration. It lands the day before AMD’s AI event, July 22 and 23, where I will be in the room both days.

NVIDIA sized an entire CPU around how a model behaves between GPU calls.

You Cannot Throw Cores at a Sequential Loop

Cloud economics bought 9x more cores since 2014 and 2x per-core speed, because chiplets made cores cheap. The agent loop only spends the 2x.

The cloud era bought 9x cores and 2x per-core speed. The agent loop only spends the 2x.

The slide that carried the argument: a trace of ten coding agents deployed inside NVIDIA. Green bars mark time with the model on the GPU; gray bars mark time acting on the CPU — tool calls, SQL queries, API calls, scripting. An agent session runs that loop 100 to 300 times, each pass through the gray sequential, the GPU waiting before the next reasoning step. NVIDIA drew the design consequence: “You cannot simply throw more cores at that to make it go faster. You actually need a more performant core.”

One caveat before the argument runs away: the trace does not prove that every gray second is CPU-bound. API calls, storage requests, and database operations can also wait on networks or remote services, and no core shortens a round trip to somebody else’s datacenter. Vera earns its premium only where its core, cache, memory, and fabric actually compress that interval.

Where it does, the metric shifts underneath. Tokens per second prices the green bars only. An agent task is billed end to end, and the most expensive meter in the datacenter keeps running while the CPU works: on the critical path, every second Vera can remove from the gray is a second Rubin no longer spends waiting at GPU rates. The unit of account for the agentic era is cost per completed task — the Denominator Stack from last week’s harness piece: “Sellers converged on the denominator, and the buyers were already measuring in it.” Vera is the silicon floor of that stack, and NVIDIA’s application benchmarks are task metrics rather than token metrics — RL evaluations completed in a window, sandbox jobs finished, and P99 latency on a trading feed. Compress the gray and cost per task falls with tokens per second unchanged.

So NVIDIA built the core. Vera carries 88 Olympus cores, custom-designed by NVIDIA for data-center workloads rather than using the licensed Arm core found in Grace: a wider, faster core tuned for the branchy, pointer-chasing code agents generate, sharing a 164MB system-level cache on one monolithic compute die, fed by SOCAMM2 LPDDR5X at up to 1.2 TB/s per socket. Pressed on manufacturability, Ian Finder, NVIDIA’s head of datacenter CPUs, scoped the monolithic claim with a candor you rarely get from a chip vendor: “We actually are using chiplets. We use them for our memory controller die... we have an IO die.” So “monolithic” means the compute die, and the testable claim is that latency under load stays flat where a chiplet-fabric CPU saturates.

88 custom Olympus cores and a 164MB system-level cache share one compute die. The fabric is what NVIDIA is actually selling.

176 vs. 256: The Claim That Matters

AMD EPYC Turin 9755. Two sockets, 256 cores, 512 copies. GCC 15.2, identical optimization flags, Micron DDR5-6400.

That is the whitepaper’s test-configuration fine print. Chip vendors benchmark against unnamed “leading x86” silhouettes as a rule; NVIDIA named the part, listed the compiler, matched the flags down to the jemalloc allocator, and published the copy counts on both sides.

SPECrate 2026_int_base — a widely followed measure of aggregate integer throughput — estimated, dual socket against dual socket: Vera 925, EPYC Turin 9755 898. That is 176 cores outscoring 256 — 3% more score from 31% fewer cores. The caveats ride with the number: vendor-run, estimated on a non-SPEC-compliant reference system, and while the SPEC suite itself is fixed, NVIDIA’s exhibit selection and commentary emphasize the subtests closest to agentic work. Anyone with a Turin box can attempt the AMD side tomorrow, and Phoronix has already tested Vera independently — although NVIDIA disabled power and frequency monitoring and limited the initial workload list to Vera’s target domains. Independent, on NVIDIA’s terms. The posture I set in May in The Last x86 Island holds: “Treat the performance as measured. Treat the efficiency as asserted.”

176 cores score 925, 256 cores score 898, vendor-run and estimated. Fleet throughput was the last x86 defense inside the agentic segment.

The 925 matters more than any per-core claim because it attacks x86’s strongest remaining defense. That defense is one I would argue myself: the agent loop is sequential per session, but a datacenter runs thousands of sessions at once, and fleet-level concurrency is a throughput problem, which is what core count buys. The socket number meets it on its own terms — if 176 cores ship more aggregate work than 256, the buyer is paying for 80 extra cores and the power to feed them. NVIDIA now claims all three tiers at once: per-core speed, socket throughput, and memory bandwidth per core.

In NVIDIA’s own loaded-latency testing, Turin falls off a cliff between 350 and 400 GB/s of delivered memory bandwidth while Vera stays flat past 1 TB/s, with a claimed 40% lower peak loaded latency and 12.7 GB/s of memory bandwidth per core against Turin’s 3.1. One caveat NVIDIA does not surface: the two systems ran different memory generations — Vera on LPDDR5X-9600, Turin on DDR5-6400 — and Turin’s cliff sits roughly where a DDR5-6400 subsystem runs out of headroom. NVIDIA chose not to run a same-memory comparison, and AMD has an obvious rebuttal available on Wednesday, July 22: NVIDIA compared two different memory generations and may be attributing too much of the resulting gap to architecture — a memory gap AMD can close far faster than an architecture gap. Caveat priced, the accounting holds — the system-level ledger Rick Xie and I ran in The Bandwidth Tax, now applied to the x86 socket with the line item named.

NVIDIA’s read: the cliff near 400 GB/s is the chiplet fabric saturating. Note the configs — Vera on LPDDR5X-9600, Turin on DDR5-6400 — so memory generation is in this gap too.

One more flag before the divider: the most serious long-run bear is not a benchmark dispute — it is software compressing the gray bars themselves. The full argument is below.

Below the divider, the stock implications: the ratio math under NVIDIA’s $200 billion claim, the memory supplier in the fine print, the honest answers on AMD and on Intel’s CPU-demand rally, the ranked winners, five bears, and the watchlist.

User's avatar

Continue reading this post for free, courtesy of Ben Pouladian.

Or purchase a paid subscription.
© 2026 Ben Pouladian · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture