The question I hear from CFOs everywhere is simple: how do we get more value from our AI spend?
For years, the market measured the success of software through adoption: seats purchased, users active, licenses renewed. Understanding the value of AI demands a more powerful measure: work accomplished.
The basic economic question facing CFOs and other business leaders is whether the value of the work AI completes grows faster than the cost of producing it.
Answering that question requires looking more deeply than a metric such as cost per token. A lower-cost model may have cheaper tokens, but getting great results may require more attempts, more time, or more human review. A more capable model may have more expensive tokens, but complete the same task in one pass. What matters is the full cost of producing a successful outcome, measured against the value that outcome creates.
The ultimate scorecard for the age of AI could be looked at as “Useful Intelligence per Dollar.” This metric answers four key questions:
- Is AI completing work that matters?
- What does each successful task cost?
- Can people depend on the result?
- Does each AI dollar produce more value as usage grows?
Start with the work itself.
How many customer issues did AI help resolve? How many code changes did it help ship? How many contracts did it review? How much time did it give back to people? How many decisions improved because the right context was available at the right moment?
Tokens create value when they transform into work people can use. As models become more capable, they can take on longer and more complex tasks: maintaining context, reasoning through multiple steps, working across tools, and adapting as they go.
The best place to begin is with one workflow. Define what “done” means and measure that outcome in the system where the work happens.
For a support team, “done” might mean a customer issue resolved. For an engineering team, it might mean a code change that passes its tests. For a legal team, it might mean a contract reviewed accurately and on time.
Consider a finance team preparing for a forecast review. Much of the work happens before a final decision is made: finding the latest forecast, moving data into Excel or Sheets, identifying changes, reconciling tabs, rebuilding slides, and checking that everything adds up perfectly.
ChatGPT Work can take on much of that process, giving the team more time to focus on the questions that matter: What changed? Why? What should we do next?
That is useful intelligence per dollar in practice. More work gets completed, faster, while people spend more of their time applying judgment, creativity, and expertise.
The next question is what it costs to complete that work well.
AI tasks vary widely. A quick answer may require little compute. A coding, research, or financial workflow may involve deeper reasoning, tool use, and many actions. Those more complex tasks can require more compute, but they can create much more value.
At the model level, cost per successful task depends on price, the amount of compute used, and the likelihood of reaching the right result. For a business, the full cost also includes employee time, human review, retries, and rework.
The calculation is straightforward:
- Add the full cost of completing the work.
- Count the tasks that met the required quality bar.
- Divide the full cost by the number of successful tasks.
This is why the lowest price per token does not always produce the lowest cost per outcome. A frontier model may deliver the best value even for a routine request if it produces the right answer in one pass, reducing retries, latency, review, and total compute.
A tiered model family gives customers more ways to optimize this equation. GPT‑5.6, which we released last week, has three tiers: Sol is our flagship; Terra balances performance and cost; Luna is our fastest and most affordable model.
These tiers provide useful starting points. The economics of the full task should ultimately determine the right model. A customer might use Luna for a fast, high-volume workflow, Terra for work requiring greater depth, or Sol when stronger reasoning delivers the best result with fewer attempts.
We trained GPT‑5.6 to get more useful work from every token. On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning set a new state of the art while using 54% fewer output tokens than another leading model. The chart below illustrates the comparison.
Across the GPT‑5.6 family, the goal is the same: more successful work per dollar. Greater efficiency makes existing tasks more affordable. Greater capability makes entirely new kinds of work possible.
Each new model generation should improve both sides of that equation. Customers should be able to accomplish more valuable work while, at the same time, the cost of completing each task continues to fall.
The third measure is dependability.
AI adoption tends to deepen in stages. First, AI helps draft. Then it finds context and reasons across tools and data. Over time, it begins taking action, handling exceptions, and completing workflows, with people providing judgment and control where needed.
Each step creates more value and asks more of the system.
Dependability has direct economic value. When results are accurate, well-sourced, consistent, and escalated appropriately, people spend less time reviewing, correcting, and repeating the work. Successful tasks cost less, and organizations gain the confidence to use AI in more important workflows.
Teams can make this concrete by tracking three outcomes:
- Ready to use: The result met the quality bar as delivered.
- Needs correction: The result required another attempt or human edits.
- Needs escalation: A person needed to step in and finish the work.
These measures tell a richer story than model accuracy alone. They show whether AI is genuinely reducing the work involved in completing the project.
Dependability also requires clear boundaries. Before AI moves from drafting to taking action, organizations should define:
- What data the system can access.
- What systems it can use or change.
- When a person should review or approve an action.
Safety, security, privacy, and control create the foundation for deeper use. People need to understand how the system behaves, how their data is handled, and how its actions are governed.
ChatGPT Work builds on the security, privacy, compliance, and workspace-management foundation of ChatGPT Enterprise. This allows organizations to give AI more context and access to more valuable workflows while maintaining appropriate oversight.
Capability earns first use. Dependability makes AI part of how work gets done.
The final question is whether the economics improve at scale.
Companies can measure this by following the same workflow over time. Track how many tasks met the quality bar, the total cost of completing them, and the cost per successful task. If completed work grows faster than total cost while quality holds or improves, each AI dollar is producing more value.
Compute sits at the center of this equation.
Compute powers research and every task that AI completes. It shapes product quality, speed, dependability, availability, and cost. Training compute builds future capability. Inference compute delivers useful work today. Both should translate into better outcomes for customers.
Better models, more efficient inference, purpose-built hardware, higher utilization, smarter routing, and stronger product design all improve the return on compute. Each generation of infrastructure helps train more capable models. Better algorithms, hardware, and software then help serve those models more efficiently.
Customers experience those improvements in human terms: better answers, faster results, fewer corrections, more dependable products, and a lower cost for the work they need done.
The gains compound. Better infrastructure accelerates research. Research produces more capable and efficient models. Better models improve products. Better products drive adoption, learning, and revenue. That growth supports continued investment in the next generation of research, compute, deployment, and safety.
OpenAI brings these pieces together through one shared intelligence platform. People use it through ChatGPT and ChatGPT Work. Developers build with it through Codex and the API. Enterprises deploy it into the systems where work happens.
When one layer improves, every product and customer can benefit.
Taken together, these four measures tell us whether useful intelligence per dollar is improving.
Useful work tells us what AI produces. Cost per successful task tells us what it takes to reach the outcome. Dependability tells us how much of the work people can confidently use. Value at scale tells us whether each dollar, and each unit of compute, accomplish more over time.
The goal is AI that helps people do more meaningful work, make better decisions, and spend more time on the parts of their jobs that require distinctly human judgment and creativity.
Our job is to make that equation better with every generation: more capable models, faster and more dependable results, and lower costs for the work customers need done.
That is how AI becomes more useful to more people and organizations over time.
Facts Only
* The economic question is whether the value of AI work grows faster than its production cost.
* Success requires measuring work accomplished rather than simple software adoption metrics.
* Tokens create value only when they transform into usable work for people.
* Cost per successful task is calculated by adding the full cost of completing the work, counting successful tasks, and dividing the total cost by successful tasks.
* The lowest cost per token does not always equate to the lowest cost per outcome.
* A tiered model family (e.g., GPT-5.6: Sol, Terra, Luna) provides options for optimizing the cost-value equation.
* Dependability is measured by outcomes: Ready to use, Needs correction, or Needs escalation.
* Dependability requires defining boundaries regarding data access, system usage, and human review.
* The goal is to ensure that completing work, faster, while people apply judgment, results in more value per dollar at scale.
* Compute powers research, task completion, and inference, and improvements across these layers compound gains.
Executive Summary
The economic value of AI spending requires shifting measurement from simple adoption metrics to measuring work accomplished relative to cost. The central question for business leaders is whether the value generated by AI work grows faster than its production cost. This involves assessing "Useful Intelligence per Dollar" across four dimensions: whether the work matters, the cost of each successful task, the dependability of the results, and whether value increases with usage.
The process starts by focusing on tangible outcomes, such as resolved customer issues or shipped code changes, rather than token usage alone. The cost calculation for an outcome must account for the full process, including human review and rework, not just the input tokens. A tiered model family, like GPT-5.6's Sol, Terra, and Luna, allows customers to optimize this equation based on the required performance versus cost for specific workflows.
Dependability is a third critical measure, focusing on outcomes like whether a result was ready to use, needed correction, or required escalation, which demonstrates how effectively AI reduces necessary human effort. This dependability requires defining clear boundaries for access and control over systems and data. Ultimately, achieving greater value at scale depends on infrastructure improvements, where advancements in compute, models, and algorithms compound the gains, leading to more capable yet cheaper outcomes.
Full Take
The narrative constructs a rigorous framework for shifting AI evaluation from superficial input metrics to tangible output economics, establishing "Useful Intelligence per Dollar" as the necessary scorecard. This framing shifts the focus from model capability (token cost) to systemic effectiveness (outcome cost and dependability). The implicit assumption is that utility derives from actionable work, not just linguistic processing.
The progression from tracking outputs (work done), costs (per task), reliability (dependability), and scaling economics reveals a complex system where infrastructure, models, and workflow design are interdependent levers for value creation. The movement toward defining dependability through the tripartite outcome structure—Ready to use, Needs correction, Needs escalation—suggests that trust is an operational metric as much as a feature.
The core tension lies in balancing capability with control. While greater model capability offers potential for entirely new work, the framework simultaneously insists that this capability must be married to demonstrable dependability and cost efficiency. The implication is that future AI development must prioritize the integration of safety, security, and governance directly into the measurement of success, rather than treating them as external constraints. The path forward demands that innovation in model design (capability) must be inherently coupled with innovation in workflow design (dependability), ensuring that economic gains are not achieved at the expense of human oversight or systemic risk.
Bridge Questions: If dependability is treated as an input variable, how should organizations dynamically adjust their compute allocation based on predicted risk profiles rather than fixed efficiency targets? What new metrics are necessary to quantify the value derived from reduced cognitive load versus the cost of error correction across diverse expert domains? How can the concept of "useful intelligence" be mathematically operationalized when measuring highly subjective outcomes like creativity or strategic insight alongside concrete task completion?
Sentinel — Human
The text reads as a high-level executive argument designed to frame the economics of AI using novel metrics, exhibiting strong conceptual coherence but lacking the mechanical uniformity often associated with pure synthetic generation.
