NVIDIA argues that AI infrastructure should be designed and built as a complete factory rather than as a set of individual accelerators. According to the post, AI factories operate continuously to produce intelligence at scale, and their economics are measured by delivered output — metrics such as tokens per second, tokens per watt, cost per token, utilization, and uptime.

The piece frames custom XPUs as a consideration for hyperscalers and AI-native companies that are building their own accelerators, situating those chips within the broader context of full-factory infrastructure.

Why it matters

Treating output metrics as the defining measure of AI infrastructure economics shifts the focus from raw accelerator specifications toward system-level efficiency and reliability. This framing bears on how large-scale AI deployments are evaluated.

Who should care

Hyperscalers and AI-native companies developing custom XPUs are the primary audience identified in the post, along with organizations planning large-scale AI infrastructure.