AI-generated article. This article was researched and drafted using AI tools and published automatically, and its featured image was generated by AI. Facts are drawn from the sources cited in the text.
Something strange is showing up on software invoices this year. The price of the underlying technology has collapsed, and yet AI costs for small business keep drifting upward anyway.
Both things are true at the same time. Neither one cancels out the other.
The price of intelligence really is falling
Numbers first, because this part is not in dispute. UBS tracks an LLM Token Expenditure Index, and it recorded the average cost of a million tokens dropping from roughly $2 at the end of May 2026 to about $1.16 in early August, a fall of around 43% in ten weeks. The South China Morning Post put enterprise AI costs at a yearly low, crediting a global price war and the fast uptake of low-cost open-source models from Chinese labs such as DeepSeek.
Zoom out further and it gets harder to believe. Epoch AI measures a median decline of roughly 50x per year in the cost of reaching a given capability level, rising to about 200x per year when the data is restricted to January 2024 onward. Google launched Gemini 3.8 Flash on September 8 at $0.75 per million input tokens and $3.75 per million output, with batch and flex inference halving that again. OpenAI cut pricing on its GPT-5.6 Luna models by 80% in late July.
For an owner who looked at these tools eighteen months ago, ran the numbers, and decided they were too expensive, that decision is worth reopening. The math has genuinely moved.
So why do the bills keep climbing?
Because cheaper units invite more units. UBS calls the effect an AI flywheel: cheaper AI encourages more usage, more usage drives demand for compute, and that demand pulls in more investment. On OpenRouter, token usage has risen roughly tenfold since the start of the year, and average monthly spending among its top 1% of users climbed from $2,500 to $7,500.
The second reason is quieter, and it catches people out. Reasoning models generate internal thinking tokens before they answer. You pay for those tokens. You never see them. Analyses from Monte Carlo and Optimum Partners both put the effect at a task costing $1 on a standard model running $5 to $20 on a reasoning model. Add agents that call themselves in loops, retrieval that stuffs long documents into every prompt, and multimodal work on images and audio, and consumption climbs faster than price falls. Navya AI’s 2026 cost report frames it bluntly: token prices down nearly 99%, enterprise bills tripled.
What AI costs for small business actually look like
Most small businesses are not buying tokens directly. They are buying seats, a per-user subscription with the token cost folded invisibly inside it. That hides the volatility, which is comfortable. It also hides where the money goes, which is not.
The more useful question is not what a tool costs per month. It is what the tool is replacing, and by how much. A seat at $30 that saves four hours is cheap. The same seat at $30 that gets opened twice is not, and vendors are quietly counting on nobody checking. That is the same discipline behind a build versus buy decision, and the same one that separates the projects that reach production from the ones that stall in pilot.
The floor is collapsing, the ceiling is rising
Worth naming, because the headlines only report half of it. Axis Intelligence describes an inference market that has split in two: the cheap end is collapsing while frontier model pricing has risen roughly 100% since January 2026, as each new generation commands more. Cheap AI and expensive AI are both getting more extreme. For most ordinary business work the cheap end is now good enough, and smaller models often fit the job better than frontier ones anyway.
One small thing this week
Pull the last three months of AI-related invoices and write two columns. What was paid. What changed in the business because of it. Not a full audit. Twenty minutes.
Most people paying for these tools are still not using them anywhere near their potential, and a falling price does nothing about that. A cheap tool that sits unopened is still waste, just cheaper waste. The genuinely interesting thing about this year is not that intelligence got cheap. It is that getting value out of it still comes down to whether anyone decided what it was for.