OpenAI Builds a Full-Stack AI Strategy Around Custom Chips, Compute and Efficiency
OpenAI says its AI strategy increasingly depends on optimising the entire stack, from custom inference chips and data centres to models, software and products, with efficiency becoming a key economic advantage.
Xcademia Team
Xcademia Research Team

OpenAI is increasingly positioning artificial intelligence as a full-stack infrastructure challenge rather than a model-only race.
In a company post published on August 25, 2026, OpenAI CFO Sarah Friar described an integrated strategy spanning data centres, chips, frontier models, developer infrastructure, consumer and enterprise products, and AI-native devices. The company argues that improvements across these layers can reinforce one another and compound over time.
A major part of that strategy is now becoming visible through Jalapeño, OpenAI's first custom inference chip. OpenAI has published its first measured performance results for the chip, positioning it as an additional first-party silicon option alongside the commercial accelerators used across its infrastructure.
Jalapeño marks OpenAI's move into custom inference silicon
OpenAI says Jalapeño was evaluated using InferenceX, a public benchmark, with GPT-OSS 120B. According to the company, the chip delivered higher peak throughput per kilowatt and lower token latency than the commercial systems included in the comparison.
OpenAI also reported results on DeepSeek R1 and Kimi K2, saying the performance gains were not limited to a single model family.
The company's published comparison reports the following mixed-token throughput results:
GPT-OSS: 22,935 mixed tokens per second per kilowatt for Jalapeño versus 427 for the existing best system, according to OpenAI.
DeepSeek R1: 12,258 versus 118.
Kimi K2.5: 6,744 versus 120.
OpenAI presents these figures as ratios of 53.7x, 104.3x and 56.1x, respectively. These are company-reported benchmark results and should be understood within the conditions of the InferenceX comparison rather than as a universal measure of chip performance.

Why OpenAI wants control over the inference stack
The importance of Jalapeño goes beyond the individual benchmark numbers.
OpenAI says developing the model, serving software, chip, memory and network together gives it greater control over how models are deployed and the economics of serving them. The company sees tighter integration as a way to work on throughput, latency, energy efficiency and cost at the system level.
This approach gives OpenAI a first-party silicon path without abandoning external hardware suppliers.
The company says future generations of its custom silicon are already underway, although it did not provide specific details about those future chips in the announcement.
OpenAI is not betting on a single hardware supplier
The company's strategy is also notable because custom silicon is being presented as part of a broader hardware and infrastructure portfolio rather than a replacement for every external accelerator.
OpenAI says Microsoft and NVIDIA have been foundational to its growth. Its current portfolio also includes AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank. According to OpenAI, these organisations contribute different capabilities across cloud infrastructure, accelerated computing, inference, data-centre development and energy delivery.
OpenAI describes this approach through the idea of staying on the Pareto frontier, seeking an effective balance between capability, speed, reliability, efficiency and cost for different workloads.
That matters because AI workloads are not identical.
Frontier model training can place different demands on infrastructure than high-volume inference. Always-on AI agents can introduce another set of requirements around latency, power and reliability.
Rather than selecting one universal architecture, OpenAI says it wants the flexibility to match workloads with the systems that provide the strongest economics and performance for that particular use case.

Data centres become part of the optimisation strategy
OpenAI also identifies data-centre infrastructure as another area where greater control can create leverage.
The company points to Project Camellia in Georgia as an example of designing facilities around customer workloads. OpenAI says the project includes commitments involving jobs, local businesses, project infrastructure and energy costs, water conservation through a closed-loop system, and an annual independent public audit of its commitments.
The announcement does not provide a detailed technical breakdown of the facility's architecture.
Efficiency is becoming an economic metric
OpenAI's argument ultimately moves beyond infrastructure performance.
The company says the value of its integrated system should be measured by how much useful intelligence can be produced from each unit of compute. That includes improvements in models, routing, context management, software and hardware.
One example cited by OpenAI comes from the Artificial Analysis Coding Agent Index. OpenAI says GPT-5.6 Sol with max reasoning reached a new high while using 54% fewer output tokens than another leading model.
For AI customers, OpenAI connects this type of efficiency with faster results, fewer retries, longer agent workflows and a lower total cost for successful work.
The broader point is that token efficiency is not simply a model benchmark concern. If an AI system can accomplish useful work with less computation, that can affect the economics of running the system at scale.

OpenAI invokes Jevons paradox
The company also connects AI efficiency with Jevons paradox, an economic concept describing situations where improvements in efficiency can increase overall consumption because a resource becomes more economical to use.
Applied to AI, OpenAI argues that as useful intelligence becomes more capable and affordable, organisations may find more tasks economically practical.
The company gives examples such as providing tailored analysis to customers, reviewing contracts, running live financial scenarios and helping engineers test more ideas.
This is an important distinction. Efficiency does not necessarily mean that total demand for compute will fall.
If the cost of useful AI work declines sufficiently, organisations may simply use AI for more tasks.
The compounding model
OpenAI's broader thesis is that AI progress can become self-reinforcing when improvements happen across the entire technology stack.
Better hardware can make software more productive. Better software can make hardware more useful. More capable models can enable better products. Better products can increase usage and generate additional signals for improvement.
OpenAI describes this as a compounding advantage, where better technology produces better economics, which can support further investment in research, infrastructure and safety.
The strategy therefore has several connected components:
Custom silicon provides greater control over specific workloads.
A diversified infrastructure portfolio gives OpenAI flexibility across hardware and cloud providers.
Software and model optimisation can reduce wasted computation.
Data-centre development provides another layer of infrastructure control.
Product growth creates demand for increasingly efficient AI systems.
Together, these elements form the full-stack strategy outlined by OpenAI.
What this means for the AI infrastructure market
The announcement highlights a broader industry shift toward vertical integration in AI infrastructure.
As AI workloads become more specialised, companies may have stronger incentives to optimise not just models, but also the chips, memory systems, networking, serving software and data centres underneath them.
OpenAI's approach does not eliminate the need for external hardware providers. Instead, the company is describing a model where first-party and third-party infrastructure coexist, with workload requirements determining which systems are used.
For enterprises, this could mean that future AI infrastructure decisions will increasingly involve evaluating the complete cost and performance of an AI workload rather than comparing individual chips or models in isolation.
The development also reflects growing demand for useful intelligence per dollar as AI deployment moves from experimentation toward larger-scale production.
A full-stack race, not just a model race
OpenAI's latest announcement signals that competition in AI infrastructure is expanding beyond model capability.
Jalapeño gives OpenAI a first-party inference chip and a new degree of control over part of its compute stack. Its wider infrastructure strategy combines that internal capability with a broad group of external technology and infrastructure partners.
The company is ultimately arguing that the strongest AI economics emerge when hardware, software, models and products improve together.
Whether that strategy delivers a sustained advantage will depend on how effectively those layers continue to improve and how OpenAI balances first-party development with its external infrastructure ecosystem.
For now, the company has made its direction clear: AI infrastructure is becoming a full-stack optimisation problem, and compute efficiency is increasingly central to the economics of useful intelligence.
Source: OpenAI
About the Author