OpenAI’s Jalapeño chip, built for fast inference at scale, was evaluated using SemiAnalysis’ InferenceX benchmark. According to the reported results, the chip registered more tokens per user and greater throughput per kilowatt than the currently available state-of-the-art hardware.

Why it matters

Inference performance and energy efficiency are central concerns as AI systems scale. Higher throughput per kilowatt points to potential efficiency gains, while more tokens per user relates to responsiveness for individual workloads.

Who should care

Organizations running large-scale inference and those tracking AI hardware developments may find the benchmark results relevant to their infrastructure decisions.