NextFin News - OpenAI chief executive Sam Altman said the company will keep driving down the price of artificial intelligence, a pledge made as the San Francisco developer unveiled Dots, its new always-on personal agents, at its annual developer conference on Tuesday. The promise is the latest escalation in a price war that has already pushed the average cost of running an AI model to its lowest level of the year.
Altman framed the cuts as the natural result of improving technology rather than a scramble for share, telling an interviewer that OpenAI wants to keep pushing prices lower even as it scales. The remarks landed on the same day the company introduced Dots - persistent agents that work on a user's behalf across Slack, Microsoft Teams and, soon, text messages - and tightened allowances on its $200 monthly subscription plan while adding a $500 tier above it.
The combination is the story. OpenAI is betting that cheaper intelligence will unlock far larger usage volumes, and that the real revenue growth sits not in charging more per token but in embedding AI agents into daily work. Whether that math holds is the central question for a company whose annualized revenue run rate recently broke past $40 billion while its losses continue to mount.
The Price War Has a Scoreboard
OpenAI's pricing ladder has moved repeatedly in 2026, and the pace has accelerated. On July 30, the company cut the price of its fastest, cheapest tier by 80%, taking it from $1 per million input tokens and $6 per million output tokens down to 20 cents and $1.20. The mid-range tier fell 20%, to $2 and $12 per million tokens. In late August, the flagship short-context model was reduced by roughly a fifth, to $4 per million input tokens and $20 per million output tokens, from $5 and $30.
Then, a week before DevDay, OpenAI replaced that entire family. Its new GPT-6 Sol and GPT-6 Luna tiers launched at $2 and $10 per million tokens, and 10 cents and 50 cents, respectively - half the promotional price of the GPT-5.6 tiers they replaced. The flagship GPT-6 Astra sits at $10 and $50 per million tokens. In under three months, the price of a unit of OpenAI frontier intelligence has fallen by more than half at the budget end of the lineup.
The moves are not isolated. An index tracking API and open-weight inference pricing maintained by a US research firm showed the average price per million tokens fell to between $1.16 and $1.18 in the first week of August, the lowest reading of 2026. That was down from $2.04 at the end of May and $1.45 in late July - a decline of roughly 43% in just over two months. A global investment bank, citing the same data, attributed the drop to a heated price war among frontier labs and a surge in adoption of low-cost Chinese open-weight models.
Competitors are moving in the same direction. Anthropic's newest flagship model is listed at $5 per million input tokens and $25 per million output, a third of the rate carried by the prior generation's top tier two releases earlier. Google and a group of Chinese open-weight model builders, including DeepSeek, have pushed budget tiers lower still, with some flash models priced in the single-digit dollars per million tokens. The result is a market in which frontier capability is getting cheaper faster than at almost any point in the software era.
"We will continue to keep dropping prices dramatically," Altman said in a separate interview earlier this month, adding that the company remained confident in the trajectory of the technology.
The mechanism behind the decline has three parts. First, each generation of model delivers more capability per unit of compute, so the same answer costs less to produce. Second, specialized inference chips and software optimizations - better batching, prompt caching, and mixture-of-experts routing that activates only the relevant sub-networks of a model - squeeze more tokens from the same hardware. Third, competition among a handful of well-funded frontier labs forces the savings through to customers rather than letting them sit as margin.
That third channel is what makes this cycle different from a normal product price cut. In most software markets, a vendor lowers price to clear inventory or win a deal. Here, the unit cost of intelligence is falling structurally, and every lab faces the same choice: pass the savings on and grow volume, or hold price and watch workloads route to a cheaper rival. The price war is not a tactic. It is the market's way of discovering what intelligence is worth when scarcity ends.
Why Cheaper Tokens Are Not the Whole Story
The same day Altman pledged lower prices, OpenAI reduced what the $200 monthly Pro plan buys. Starting October 30, the allowance for its ChatGPT Work and Codex applications falls from 20 times the Plus plan's allowance to 10 times, and the weekly message cap on its top consumer model drops from 200 to 100. At the same time, the company introduced a $500 monthly Pro tier with the highest usage limits and an "Ultrafast" speed mode that generates up to 300 tokens per second.
That juxtaposition is the key to reading the strategy. OpenAI is cutting the unit price of raw intelligence while reserving the highest-volume, lowest-latency access for the customers willing to pay the most. It is a two-tier playbook with a clear logic: commoditize the input, monetize the throughput.
Dots sits at the center of that design. The agents are included with eligible Pro and Business Premium subscriptions, and conversations with a Dot are reported not to count against a user's message limits. In effect, OpenAI is making the agent interface free at the margin while the subscription tiers become the gate. The bet is that users will not pay for tokens - they will pay for an assistant that never stops working.
The financial logic behind the price cuts is equally clear, and it is unforgiving. OpenAI's annualized revenue run rate held near $25 billion from February through May before breaking higher, reaching roughly $40 billion by August, according to reporting on the company's internal metrics. The chief financial officer told investors in mid-August that enterprise revenue had overtaken consumer revenue for the first time, crossing a line the company had previously expected to reach only at the end of the year; enterprise had stood at just over 40% of revenue at the time of the company's March funding round and has since moved past half. Business revenue rose 32% in a single month.
Yet the cost side is rising just as fast. The company posted a 33% gross margin in 2025, and its inference bill - the compute cost of running models - was about $8.4 billion last year, with projections pointing to roughly $14.1 billion in 2026. Full-year 2025 revenue was $13.07 billion against $34 billion in costs and expenses, an operating loss of about $20.9 billion, and internal projections point to a loss of roughly $14 billion in 2026, with profitability not expected until the end of the decade.
For the price cuts to expand rather than compress margins, usage must grow faster than the per-token price falls. That is the volume hypothesis: a tenfold drop in price must produce more than a tenfold rise in demand. History in software offers some support - cloud computing followed a similar path, and the advertising business that runs on cheap attention is proof that near-zero marginal cost can underwrite enormous scale. But the timing is unforgiving when the compute bill is doubling and the loss column is already larger than revenue.
Cyclical Pressure or a Structural Reset?
The right way to read this moment is as a structural reset with a cyclical overlay, and the distinction decides the investment conclusion. The structural leg is durable: model efficiency improves every generation, specialized silicon lowers the cost per inference, and competition among a small group of capital-rich frontier labs forces savings through to customers. None of those forces self-correct. A model that needs half the compute to produce the same answer does not suddenly need more.
The cost history supports the structural read. The price of GPT-4-class capability has fallen roughly tenfold since 2023, from about $30 per million input tokens to the low single digits on budget tiers today. That decline is being driven by algorithmic progress that compounds - better architectures, better training data, better inference-time routing - rather than by one-off discounts. Algorithmic efficiency, isolated from hardware price declines and competition effects, has been estimated to improve at a steady pace year over year.
The cyclical leg runs the other way. Demand for AI is surging - enterprise deployments, agent workloads, and consumer usage are all climbing - and the industry is in the middle of a capital-spending cycle to build the data centers and power infrastructure to meet it. When demand outpaces supply, prices stop falling. There are already signs of friction: the blended average price index can move on mix-shift, as more workloads route to cheaper models, which makes the apparent price decline look steeper than the underlying cost decline. And capacity additions are not frictionless - chips, power contracts, and cooling all take time to bring online.
Separating the two legs matters because they point to different conclusions. If the structural leg dominates, AI becomes a cheap, ubiquitous input - a utility - and the winners are the platforms that own the distribution layer where users and agents meet. If the cyclical leg dominates in the near term, the price war pauses, margins stabilize, and the companies that locked in long-term compute supply gain an advantage over rivals buying at spot rates.
Over a full investment horizon, the evidence favors the structural read. But over the next few quarters, the cyclical leg is likely to produce bumps: capacity constraints, power limits, and the sheer pace of agent adoption can all slow the descent. The practical implication is that the price curve will slope down over years but not in a straight line over months.
The Counter-Case: Price Cuts as a Sign of Strain
The strongest argument against reading the cuts as pure confidence is that they are also a defense. OpenAI faces competition on three fronts: Anthropic, which has grown its revenue faster in recent months and leads in some developer mindshare; Google, which bundles AI into products billions of people already use and can afford to subsidize intelligence indefinitely; and open-weight models that offer good-enough performance at a fraction of the cost. In that environment, holding price can mean losing share, and losing share in a winner-take-most market can be fatal.
There is also the balance sheet. The company's 2025 operating loss exceeded its revenue, and internal projections point to a loss of roughly $14 billion in 2026, with profitability not expected until 2029 or 2030. Cutting prices while posting losses that size is only rational if it buys growth that cannot be bought any other way. Skeptics argue the company is trading margin for momentum in a market where momentum is exactly what is hardest to keep, and where the next generation of models could reset the leaderboard at any time.
The answer to that counter-case is that the alternative is worse. In a market where the unit cost of intelligence is falling structurally, a company that holds price does not preserve margin - it loses volume and then margin anyway, as customers route workloads to cheaper models. The price cut is the admission that commoditization is coming; the only choice is whether to lead it or be caught by it. The companies that tried to hold premium pricing through a technological deflation have a long history of ending up with neither margin nor share.
Two signals would prove the structural read wrong. First, if the average inference price across major providers rises for two consecutive quarters while capacity additions remain on schedule, the decline is demand-driven scarcity rather than structural progress, and the utility thesis breaks. Second, if OpenAI's gross margin does not expand as usage scales - if the volume hypothesis fails and the loss column keeps growing faster than revenue - the price cuts are buying share, not efficiency, and the path to profitability moves further out.
What Comes Next
In the short term, the price war is likely to continue but at a slower pace. The low-hanging efficiency gains have already been taken, and the next round of cuts will require the following generation of models to clear a higher bar. Watch the pricing pages of the three frontier labs over the next quarter: another 20% cut to a flagship tier would confirm the structural trend; a pause would signal that the cyclical leg is tightening.
Over the medium term, the battleground shifts from price to distribution. Dots, and agents like it from Meta, Google, and others, are the first serious attempt to own the interface where work happens. The company that becomes the default place users send their tasks - across email, documents, code, and messaging - captures the subscription revenue regardless of which model runs underneath. That is why the agent launch and the price pledge landed on the same day: they are two halves of one strategy. Cheap tokens pull demand forward; the agent owns the relationship.
In the long term, intelligence looks set to become a cheap input, and the economic surplus will migrate to the layers above it - the applications, the workflows, and the proprietary data that turn cheap tokens into valuable outcomes. The second-order consequence is the one most investors are not pricing: if intelligence is cheap for everyone, then no software company can defend a margin premium on the strength of its AI alone. The moat moves from the model to the data, the distribution, and the workflow lock-in. That is a bearish read for pure model vendors and a bullish read for the platforms that sit between the model and the user.
Scenarios for the next 12 months split cleanly. In the base case, prices drift lower by 20% to 40% annually, usage more than doubles, and gross margins inch up as efficiency gains outrun the compute bill - a difficult but navigable path. In the upside case for OpenAI, agents become the dominant interface, enterprise adoption accelerates, and the $40 billion run rate compounds into a profitable utility-scale business by the end of the decade. In the downside case, the volume hypothesis fails, the compute bill keeps doubling, and the company is forced to raise prices or dilute shareholders again to fund the build-out.
Altman's pledge is easy to state and hard to keep: prices fall only as long as technology outruns demand. The real story is not whether OpenAI wants cheaper AI - every lab does. It is whether the company can build a business large enough to survive the world it is helping to create.
Explore more exclusive insights at nextfin.ai.
