Skip to main content
OpenAI

July 31, 2026

Company

Building abundant intelligence

A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.

Loading…

AI infrastructure is not valuable because it is large. It is valuable because of what it makes possible: more capable intelligence, available to more people, at a lower cost.

That is how I think about abundance. It is both embedded in our mission—ensuring that artificial general intelligence benefits all of humanity—and in the economic engine that drives our business.

When the cost of useful intelligence falls, more work becomes worth doing. When models become more capable, that work creates more value. As adoption grows, we gain the revenue, real-world feedback, and visibility into demand to keep investing in the next generation of research and infrastructure.

Better intelligence drives broader adoption. Broader adoption supports more investment. More investment improves intelligence and efficiency. That is the cycle we are building.

The economics of abundance

Our recent pricing announcement shows how that cycle quickly benefits customers. Yesterday, we reduced the price of GPT‑5.6 Luna by 80 percent and GPT‑5.6 Terra by 20 percent. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens; Terra costs $2 and $12, respectively.

For GPT‑5.6 Sol, Fast mode delivers up to 2.5 times the speed of standard processing at twice the price, with no change in intelligence.

These are not simply changes to a price list. They expand the range of work that becomes practical and give customers more flexibility to balance intelligence, speed, reliability, and cost.

The right question is not which model belongs to which task. It is how much intelligence the outcome demands for the required result, how quickly the intelligence is needed, and what the intelligence should cost to achieve. That balance may change several times within the same workflow.

Customers do not buy tokens for their own sake. They want the support issue resolved, the software shipped, the contract reviewed, or the scientific question answered. The right measure is the cost of a successful outcome, including the time, retries, oversight, and errors required to get there.

A stronger model that completes the work correctly and efficiently can, at the end, be more economical than a cheaper model that requires repeated attempts or extensive human intervention. Conversely, a lower-cost model can dramatically expand access when it meets the same quality bar. The opportunity is to apply the maximum useful intelligence at the right price.

Getting more from every unit of compute

Delivering that value requires more than building additional data centers. It requires making every unit of compute more productive.

Our recent engineering work illustrates what that looks like. Working with our technical teams, GPT‑5.6 Sol helped optimize the production software used to serve our models, reducing end-to-end serving costs by 20 percent. It also helped improve speculative decoding, increasing token-generation efficiency by more than 15 percent.

Just as important, efficiency is not determined by the model alone. The system surrounding it matters. Better routing keeps hardware productive. Smarter context management prevents agents from repeating work. Stronger tools and product design reduce the number of steps required to complete a task.

In a recent benchmark analysis, improvements to retained reasoning and context management raised GPT‑5.6 Sol’s score on the public ARC-AGI-3 task set from 13.3 percent to 38.3 percent while using six times fewer output tokens. The model did not change. The surrounding system did.

These gains compound. More capable models help our teams discover new efficiencies. Those efficiencies lower the cost of serving customers and expand the work we can support with the same infrastructure. That, in turn, makes the next generation of intelligence more accessible.

Why the full stack matters

The advantage of building across infrastructure, models, platform, and products is not simply that we participate in each layer. It is that each layer makes the others better.

Real-world product use shows us where customers find value and where they encounter friction. That feedback helps shape our research. Research improvements strengthen our products and lower the cost of serving them. Demand across ChatGPT, ChatGPT Work, Codex, and the API helps us make better decisions about where to add capacity.

The scale of that learning matters. Our models now reach more than one billion active users and more than two million businesses. As people gain confidence in the technology, they use it more deeply. Six months after signing up, people send roughly 50 percent more messages each day and use ChatGPT for about twice as many kinds of work. ChatGPT Work is changing what it means to be a knowledge worker, moving beyond answering questions to completing complex, multistep work—“asking” to “doing.” Across OpenAI, agentic work through Codex now accounts for 99.8% of weekly output tokens, with Finance among the teams that have made agentic tools a primary part of how they work.

Inside organizations, the pattern is similar. A company may begin with one team or one workflow. As the quality and economics improve, adoption spreads across functions and AI becomes part of how the business operates.

That creates a valuable feedback loop between demand, product improvement, efficiency, and infrastructure planning. It does not require owning every asset or building every component ourselves. We can own, partner, or buy depending on what best serves the customer and makes the most economic sense. What matters is coordinating the system and learning across it.

Building with conviction and discipline

AI infrastructure must be planned years before it is needed, while models, products, and customer demand evolve much faster. That mismatch makes discipline essential.

We base our investment decisions on evidence: user and workload growth, enterprise commitments, API consumption, utilization, revenue, and progress in model capability and efficiency. Technical and commercial milestones help determine when projects advance. Long-term partnerships bring together the financing, infrastructure, and operating expertise required to deliver at scale.

Product revenue, private capital, and commercial partnerships play different roles in supporting that growth. The objective is not to build the most infrastructure. It is to deploy the right capacity, at the right time, against credible demand.

For me, the key questions are straightforward: How quickly does new capacity become productive? How efficiently is it used? What customer demand does it support? How rapidly can technical progress lower the cost of delivering useful intelligence?

Those questions connect long-term ambition to operating discipline.

The opportunity ahead

We are still early. More capable systems will complete longer projects, coordinate across tools, and handle more of the work between an idea and a finished result. Individuals and small businesses will gain capabilities once available only to much larger organizations. Enterprises will apply intelligence more broadly across their operations.

Our goal is not simply more compute, bigger models, or lower token prices. It is more useful intelligence within reach.

That is what abundance means: intelligence that keeps getting more capable, more affordable, and more valuable to the people who use it. We will measure our progress by how much useful work it makes possible, how efficiently we deliver it, and how widely its benefits can be shared.

Author

Sarah Friar