Thunder Bay AI
The Journal
TrendsAugust 26, 2026 5 min read

AI prices keep falling: what that means for a small-business budget, and when to wait

Model prices at the API layer dropped sharply in mid-2026. For most small businesses paying a monthly subscription, those drops have not passed through. Here is the distinction that matters, and how to think about timing.

AI model prices at the API layer have been falling sharply through 2026. On July 30, OpenAI cut its GPT-5.6 Luna model by 80 per cent — from $1.00 to $0.20 per million input tokens — and reduced GPT-5.6 Terra by 20 per cent; in August, Anthropic confirmed that Claude Sonnet 5 will remain at its introductory $2.00 per million input tokens indefinitely, cancelling a planned September 1 increase to $3.00. But for most small businesses in Northwestern Ontario — who access AI through monthly SaaS subscriptions like Microsoft Copilot, ChatGPT Plus, or Notion AI — those model-layer drops have not translated into lower bills. The price compression is real and will continue; whether your business benefits from it now depends on how you access AI, not just how much you pay.

Where prices have actually fallen

The cuts happening in 2026 are at the API layer — the pay-per-use pricing charged to developers and businesses that build directly on AI model infrastructure. OpenAI's GPT-5.6 Luna, after the July 30 reduction, costs $0.20 per million input tokens and $1.20 per million output tokens. Anthropic's Claude Sonnet 5 is $2.00 per million input tokens and $10.00 per million output tokens — pricing that was announced as temporary but made permanent in August 2026, verified on the Anthropic pricing page. A million tokens represents roughly 750,000 words of processed text; for a single business task on a few hundred words, the per-use API cost is a fraction of a cent. Anthropic's Batch API offers an additional 50 per cent discount on both input and output tokens for non-time-sensitive processing, and prompt caching reduces the cost of re-processing repeated content to 10 per cent of the base input rate. The underlying cost of running an AI inference has genuinely dropped.

Why your subscription price has not moved

Most small businesses do not access AI at the API layer. They access it through finished SaaS products — monthly per-seat subscriptions to tools built on top of AI models. A vendor charging $20 or $30 per user per month prices based on the value the product delivers, not the marginal cost of the model inference underneath. When underlying model costs drop, vendors absorb the improvement as margin or reinvest it in product development; they do not typically pass it through as subscription price reductions. A 2026 analysis of 2,457 active AI tools found that 65.1 per cent require direct payment and only 5.0 per cent are genuinely free — figures consistent with a market where vendors have not responded to falling model costs by reducing subscription prices. The exception tends to be tools with low switching costs and high competition, such as productivity and education products, where free and freemium options have spread. The paid SaaS tools most businesses rely on for operational or customer-facing work sit in a different pricing environment.

When it makes sense to wait

The case for waiting is sharpest when you are evaluating a large custom AI build — a project where engineering or integration costs dominate the budget and the underlying model inference cost is a meaningful ongoing operating expense. In that scenario, prices are likely to continue falling, and a measured delay may reduce long-term operating costs. The case falls apart in two situations. First, when there is an active operational problem: a task currently consuming staff time, generating errors, or losing customers. The cost of inaction accumulates daily; a model that is cheaper in six months does not recover what the delay cost. Second, when there is available funding. FedNor's Regional Artificial Intelligence Initiative covers up to 75 per cent of eligible non-capital project costs for Northern Ontario businesses adopting AI, with continuous intake — confirm current eligibility and intake status directly with FedNor before applying, as retail businesses are excluded and a FedNor officer contact is required before a formal application. When a funding program absorbs most of the implementation cost, the direction of API prices is secondary to the funding window.

The practical question for an NWO business

The useful frame is not "should I wait for prices to fall further?" but "what am I actually paying for, and what problem does it solve?" For a business paying a per-seat subscription to an AI writing, scheduling, or customer service tool: subscription prices are unlikely to drop in the near term, but the tools themselves improve as underlying models do — more capability for the same invoice amount. For a business considering a custom integration — connecting AI to a specific workflow, a proprietary data set, or a customer-facing channel — falling API prices matter more, and the timing question is whether the operational problem is urgent enough to justify implementation now versus six to twelve months from now. For businesses in capital-intensive NWO sectors — mining, forestry, healthcare — sector-targeted AI funding programs can change the cost calculus entirely. Contact FedNor at 1-877-333-6673 or the Northwestern Ontario Innovation Centre at nwoinnovation.ca to confirm current program availability before making a timing decision based on price trajectory alone.

The most important distinction when reading AI cost headlines: model-layer API pricing (what a developer pays to call an AI model) and SaaS subscription pricing (what a business pays for a finished AI product) are different markets with different pricing dynamics. A large percentage reduction in model API pricing does not predict a comparable reduction in the monthly subscription fee for a SaaS AI tool. If your AI spend is subscription-based, monitor vendor pricing pages directly — do not assume headline model price drops have passed through to your bill.

Sources: OpenAI GPT-5.6 price cuts July 30, 2026 — Terra reduced 20% (to $2.00/$12.00 per million tokens), Luna reduced 80% (to $0.20/$1.20 per million tokens): cloudzero.com/blog/openai-pricing/ | Anthropic Claude Sonnet 5 pricing made permanent at $2.00/$10.00 per million input/output tokens, cancelling planned September 1 increase to $3.00/$15.00; Batch API 50% discount; prompt caching 10% cache-read rate — verified directly: platform.claude.com/docs/en/about-claude/pricing | AI tool payment-required rate (65.1% of 2,457 tools) and genuinely free rate (5.0%) — 2026 analysis: tooldirectory.ai/blog/ai-model-prices-fell-tool-prices-didnt-2026 | FedNor Regional Artificial Intelligence Initiative — up to 75% non-capital / 50% capital (repayable), continuous intake, retail excluded: fednor.canada.ca/en/our-programs/regional-artificial-intelligence-initiative-raii-northern-ontario | Northwestern Ontario Innovation Centre: nwoinnovation.ca

Frequently asked

If AI API prices keep falling, should I wait before subscribing to an AI tool?

For subscription tools — monthly per-seat products — waiting is unlikely to result in a lower price. Subscription pricing follows the value vendors deliver, not the cost of the underlying model. Where waiting may make more sense is before commissioning a custom AI build, where model inference is an ongoing operating cost. Even then, if there is an active operational problem you are losing time or revenue on, the cost of delay accumulates and may exceed the savings from lower future prices.

What is the difference between AI API pricing and a SaaS subscription?

API pricing is what a developer pays to call an AI model directly, charged per unit of text processed (tokens). It is the infrastructure layer. SaaS subscriptions are finished products built on top of those models, charged per seat per month. When model API prices drop, the SaaS vendor benefits through lower costs; the end subscriber does not necessarily see a change in their invoice.

Does FedNor RAII funding change the timing question for AI adoption?

Yes — when an eligible funding program covers a significant share of implementation costs, the program's availability typically matters more than the direction of API prices. If RAII is open and your project qualifies, waiting for lower API prices while the funding window closes may cost more overall than acting now. Confirm current eligibility and intake status directly with FedNor at 1-877-333-6673 before applying — retail businesses are excluded and a FedNor officer contact is required before a formal application is submitted.

Share this brief

Get the weekly Signal

The AI, funding, government, and tech moves that matter for Northwestern Ontario — one email a week, source-linked, read by a human before it reaches you.

One email a week from Thunder Bay AI. Unsubscribe anytime.