Something strange is happening in the AI market. The product has never been cheaper, and the companies selling it have never lost more money. Both things are true at the same time, and if you run security for a company whose employees use these tools, you should understand why.
The price collapse is real
Start with the numbers. When the GPT-3 API launched in June 2020, it cost about $32 per million tokens. Today the benchmark sits below ten cents per million tokens. That is a drop of more than 300 times, according to IDC research reported this fall.
The cuts keep coming. On September 22, 2026, OpenAI and Anthropic shipped new models on the same afternoon and both cut prices. OpenAI launched GPT-6 Sol at $2 per million input tokens and $10 per million output tokens, down 50 percent from its predecessor, alongside GPT-6 Luna at just ten cents and fifty cents. Anthropic answered with Claude Opus 5.5 at $4 and $20, a 20 percent cut, with cache reads slashed 60 percent. In early October, Anthropic cut its Haiku model 75 percent to ten cents per million input tokens.
Per-token prices for comparable capability have fallen more than 90 percent since 2023, driven by competition, better model architectures, cheaper inference chips, and open-weight models like DeepSeek undercutting everyone.
Tokens are the unit of cost, and developers are shrinking them
A quick primer, because the whole pricing model rests on it. Language models do not read words. They read tokens, chunks of text averaging about four characters each. Every prompt you send and every word the model generates is metered in tokens. When labs quote $2 per million input tokens, that is the price of roughly 750,000 words in.
That metering is why so much engineering effort now goes into making tokens cheaper. The biggest lever is matching the model to the task. A frontier reasoning model is overkill for summarization, extraction, or classification, so developers route simple work to small, fast, cheap models like Haiku or Luna and reserve the expensive models for complex coding and agentic work. The labs themselves now sell this two-tier structure. Other levers stack on top: prompt caching reuses repeated context at a fraction of the price, distillation trains small models to mimic big ones for narrow jobs, and leaner prompts simply burn fewer tokens per answer.
This is the quiet engine behind the price collapse. It is not only that inference got more efficient. It is that the industry learned to stop spending flagship tokens on commodity work.
Agents are where the value actually is
Here is the paradox. Agents are the most expensive way to use AI and the most effective way to use it.
A chatbot answers one question per prompt. An agent takes an objective, plans the steps, calls tools, queries databases, runs code, checks its own work, and iterates until the job is done. That loop is why a single agent task can burn 96,000 tokens before producing an answer, and why the labs lose money on power users. But it is also the first form of AI that completes real work instead of drafting text about work.
For enterprises, this is the unlock. A customer service agent that resolves tickets end to end, a coding agent that ships tested features, a research agent that compiles cited briefings. Those replace labor costs measured in salaries, not token costs measured in fractions of cents. The unit economics look terrible per token and excellent per outcome.
That is why I expect the industry to break even within about two years as agents move into wide enterprise adoption. The losses today are concentrated in consumer subscriptions priced for 2023 chatbot habits. Enterprise agent deployments are priced per seat, per workflow, or per outcome, and the value side of that equation is enormous. Once agents are doing billable work at scale, the math flips. The subsidy era ends not because AI gets more expensive, but because it finally gets priced against the labor it replaces.
The losses are just as real
Now the other side. OpenAI generated roughly $20 billion in revenue in 2025 and lost about $8 billion doing it. In the first quarter of 2026, the company lost $1.22 for every dollar of revenue it brought in, an adjusted operating margin of negative 122 percent, according to reporting from The Information. It projects $14 billion in losses for 2026 alone.
The reason is usage. SemiAnalysis tested every ChatGPT subscription tier by maxing out weekly usage with coding and agentic tasks, then compared actual token consumption to public API pricing. A fully used $200-per-month ChatGPT Pro plan consumes compute equivalent to about $14,000 in API costs. That is a 70-to-1 gap. OpenAI starts losing money on Plus and Pro plans at just 11.4 percent utilization, and hits zero gross margin on the top Pro tier at 5.7 percent utilization. A developer using the tool daily for real coding work crosses into loss territory for OpenAI well before hitting any limit.
The culprit is agentic workflows. A typical AI agent job burns around 96,000 tokens before producing a single answer, up to 1,000 times more tokens than casual chat. OpenAI priced its subscriptions for 2023 chatbot habits. The product is now used for 2026 autonomous agent workflows that were never in the pricing model.
Anthropic shows the problem is solvable but not solved. Its gross margin was negative 94 percent in 2024. By 2026 it has climbed to roughly 60 percent, mostly through inference efficiency: nine months ago Anthropic generated $16 million in annual recurring revenue per megawatt of compute, a figure expected to reach $60 million. Even so, a fully used $200 Claude Max plan still costs the company about $8,000 in compute, a 40-to-1 gap. Both labs lose money on heavy users running agents at sustained intensity.
Why security teams should care
Here is the part that matters for my day job. When intelligence is nearly free, consumption explodes. Every employee becomes a power user. Every team wires agents into business processes. Shadow AI stops being a few people pasting into a chatbot and becomes thousands of autonomous workflows calling models, retrieving company data, and touching production systems.
The labs are subsidizing this adoption with investor money. Your company will absorb the risk for free. Cheap inference means more agents, more data flowing into third-party models, more non-human identities with access to real systems, and more output nobody reviewed.
The economics will eventually correct. Prices will rise, usage tiers will tighten, or the subsidies will end when the funding environment changes. The exposure your company builds during the cheap years does not correct with it. Data shared with a model is shared permanently. Agent permissions granted casually persist.
So enjoy the cheap intelligence. Just govern it like it costs what it actually costs: controlled intake for AI tools, data guardrails before the prompts, identity discipline for every agent, and logging on all of it. The labs can afford to lose money on your usage. You cannot afford to lose control of it.