Deep Water Semiconductor半导体深水区 · translated column
SemiPulse →

OpenAI Picks AMD Turin Over NVIDIA Vera for Jalapeño ASIC Hosts

OpenAI selected AMD EPYC Turin as the host CPU for its in-house Jalapeño ASIC, prioritizing deployment maturity and operational stability over the latest platform.

October 04, 2026  ·  originally in Chinese

A lab building its own chips chose a competitor's mature processor over the chipmaker's official new CPU for its first-generation inference ASIC. The details of this decision are more revealing than the silicon itself.

According to an interview published by Tom's Hardware on October 2, OpenAI Vice President and Hardware Lead Richard Ho confirmed that the in-house Jalapeño ASIC is already deployed at rack scale internally, with each host equipped with an AMD EPYC Turin CPU featuring 1.5TB of memory. This typically unglamorous component has been placed at the center of the selection debate for the first time.

To provide context, Jalapeño is OpenAI's in-house inference acceleration chip, previously known only through sporadic supply chain rumors. This marks the first official confirmation of its deployment form: SemiAnalysis described its rack-level deployment after the chip's reveal, and Tom's Hardware subsequently secured a direct response from Richard Ho. The key takeaway is less about the chip's debut and more about its deployment finalization—the chip is already running in its own racks, and the company has officially stated which host it runs alongside.

Selection Layer: Why Bypass Vera

Richard Ho cited three layers of reasoning. The first is rhythm: the project's design principle is to de-risk and deploy quickly—performance and cost targets can be aggressive, but no additional variables should be introduced to the platform. The second is a direct quote: Vera, as a standalone product, is "slightly behind in maturity"; Turin is strong and delivers all the functionality required for the project. The third is the most practical: partners have existing experience with Turin platforms, meaning the path for deployment and operations has already been paved.

Note the scope of this comparison: Ho compared the Turin and Vera products, with the conclusion limited to the Jalapeño project. When asked about the Arm versus x86 architecture debate, he did not take sides. One detail underscores the engineering considerations: Turin uses socket installation, allowing the entire host CPU to be replaced if it fails; Vera is board-level packaged, meaning a fault requires replacing the entire board. This difference, highlighted by tech media in their analysis of the selection, carries significant weight for rack clusters that must manage their own operations and maintenance.

For readers unfamiliar with the platform: Turin is the codename for AMD's fifth-generation EPYC, launched in the second half of 2024 with a focus on high core counts and memory channels. It has completed full production cycles in cloud and supercomputing environments. Vera is NVIDIA's first in-house data center CPU core, debuting with the Vera Rubin platform, and is still in the early stages of accumulating production and service experience. The assessment that it is 'slightly less mature' accurately reflects the gap between these two track records.

Deployment Layer: What 1.5TB of Memory Means

SemiAnalysis describes this deployment as 'rack-scale,' with each Turin host equipped with 1.5TB of memory. This figure warrants specific attention: in ASIC clusters, the host CPU handles scheduling, preprocessing, and inference orchestration. Large memory capacity allows for larger KV caches and request queues to be processed on the host side, reducing cross-machine access. 1.5TB represents a high-end configuration for current dual-socket servers, indicating that OpenAI is configuring the inference orchestration side for heavy loads; the host is not a peripheral component.

On October 1, we wrote about a new source of server CPU demand: every multi-step planning cycle in agentic AI requires a round of task decomposition on the CPU side (see the article from that day). That piece addressed the demand side; this one covers the deployment side of the same trend. Agentic workloads are characterized by fragmented requests, high state complexity, and dense round-trips, which are precisely the areas where host memory and host-side bandwidth are most constrained. Increasing per-host memory to 1.5TB acknowledges that orchestration loads are no longer incidental; the configuration standards for the host side are being rewritten to treat it as a primary workload.

1.5TB

Memory per host

Rack-scale

Jalapeño deployment form factor

Internal use only

Current sales plan

The three items above refer to: the memory capacity per EPYC Turin host in the Jalapeño cluster, the scale form factor of this deployment, and OpenAI's external strategy for this chip—currently restricted to internal use, with potential for broader rollout in the future not excluded.

Landscape Layer: The Host Becomes a Competitive Position

The industry implication of this selection is that the host CPU is transitioning from an accessory to an independent competitive position. NVIDIA's marketing claims that Vera offers a 1.8x performance improvement over Turin; however, technical media reviewing the test suite noted that the overall lead over the AMD EPYC 9755 is approximately 3%—the gap between promotional figures and composite scores is exactly the margin buyers scrutinize when making decisions. OpenAI remains a deployment partner for Arm's AGI CPU and is described by NVIDIA as a customer 'exploring Vera,' meaning a single selection does not constitute an architectural verdict.

For the server industry, there is an impact that never makes the press releases. The x86 host ecosystem has accumulated twenty years of operations toolchains, firmware, and channel experience. When hyperscalers build their own racks, whether a partner is familiar often matters more than benchmark scores. Ho listed this as a formal selection criterion, effectively stating the industry's unwritten rules. For a new CPU to enter these racks, it must first get past the existing operations teams. This is the truly difficult part for Vera.

My assessment is that the next-generation Jalapeño platform will not use Vera as the host unless Vera produces independent third-party benchmarks across workloads. The falsification condition is clear: if OpenAI's next-generation custom chip rack officially adopts Vera, the maturity gap has been closed, and this assessment is void.

See you in the comments. If you managed the racks, would you choose a socket-based host to save on board-level repairs, or bet on a new CPU for overall system efficiency?

This is an automated English translation of a column originally published in Chinese as《半导体深水区》. Numbers and product names are preserved from the original; wording is machine-generated and may differ from the author's intent. ← All articles