Creators' AI

Creators' AI

šŸ’µ How I stopped overpaying for AI Models

Here's how I stopped overpaying for yesterday's AI, and the tools I use to track it.

Daniil Andreev's avatar
Creators AI's avatar
Daniil Andreev and Creators AI
Oct 07, 2026
āˆ™ Paid

Hey, glad you’re here for another Creators AI edition.

I have a small request: take your time with this next line:

GPT-6 costs less than GPT-5.

I don’t mean ā€œbetter value.ā€ I mean the sticker price is lower. GPT-5.6 Sol launched in July at $5 in / $30 out per million tokens. by September, GPT-6 Sol was here at $2 / $10. Ten weeks, one lab, a smarter model, and an output price roughly a third as high.

A wall calendar with price tags getting smaller and smaller, and a happy robot with a basket

Here’s the uncomfortable bit: if you don’t keep up, you’re not merely missing some shiny new features. You could be paying more for a weaker model. In a market moving this fast, that’s a strange place to plant your flag.

And, yes, this is personal, I used to be that person. Grab a coffee, I have a mildly embarrassing story to tell :)

Here’s what we’re getting into:

  • The receipts: which prices moved, and how much

  • My confession: the old Opus I stuck with for way too long

  • A quick tour of OpenRouter

  • Cheaper Inference: tracking spend and trying cheaper models

  • A 15-minute test you can borrow

  • A starting point, whether you have a project or not

Would practical AI know-how and the news worth paying attention to be useful in your inbox?


The receipts

My memory is no pricing database, so I went back through our weekly digests instead of asking you to take my word for it. Here’s about six months of changes:

price-drops-table.png

Take a moment with the Fable 5 row. Back in July, Anthropic’s top model cost $10 / $50 (we looked at which Claude model to pick then). Today, Opus 5.5 is $4 / $20, and Anthropic says it comes close to Fable 5.1 on most work. I’d treat that as a claim, not a sacred text, but even ā€œcloseā€ would mean a very different bill.

The final row might be my favorite. GPT-5.5 doubled its price right before that, up to $30 per million output tokens. Then DeepSeek V4 Flash turned up the next day at $0.28. In the same week, that’s about 100Ɨ cheaper.

Here are a couple more quick jolts:

  • September brought DeepSeek V4.1 Flash beating GPT-5.6 Sol on four agentic benchmarks, with a price of $0.30 per million input.

  • Then, over one 72-hour stretch at the end of the month, along came Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon. You could blink and miss a whole pricing page.

One caveat before I make falling prices sound inevitable: sometimes they go up.

  • DeepSeek bumped its own prices 2.3 to 4.5Ɨ this summer, while still growing 172% on OpenRouter, which is a little funny.

  • Gemini 3.7 Flash starts at an introductory $0.75 / $3.75, then doubles to $1.50 / $7.50 on January 1.

So, yes, prices move in both directions. That’s why ā€œchoose the best model onceā€ doesn’t feel like much of a strategy anymore. It’s more like agreeing to a price and never checking it again.

Know someone who sends everything to the priciest model by default? This might be worth forwarding!

Share

My confession: the Opus I couldn’t quit

If you saw my OpenClaw migration post, you already know the setup: an always-on agent runs my agency and about half this newsletter. At one point, the bill climbed from around $25 a month to $500+ while doing the same work. The culprit was my own choice: I’d set the biggest model as the default for everything. (Been blindsided by a token bill? The token trap nobody mentioned will feel familiar.)

That was chapter one. Chapter two is ShopClaw, the agent loop behind my store. It closes tasks all day: small fixes, routine jobs, the boring stuff nobody wants to do twice.

A dusty old robot with a big price tag being carried to checkout, while a sleeker cheaper one sits ignored on the shelf

I fixed the bill once. Then I never looked again. ShopClaw ran on Opus 4.8 plus Sonnet 4.5 for more than 2 months. It worked. Nothing was on fire. Nobody was complaining. And that's the trap, because when you overpay for an old model, nothing breaks. The bill just sits there, quietly being higher than it needs to be. It's a cousin of the AI productivity tax: a cost you pay every day without ever seeing an invoice line for it.

Honestly it was my X feed which made me consider that as I saw bunch of indie hackers posting about saving on different models usage

Then Sonnet 5.5 shipped. I moved the whole loop to it, Sonnet 5.5 only, no Opus, and let it run. Then I did the thing I should have done months earlier: I pulled the telemetry. 7.5 days before the switch, 3 days after. Same kind of work. Everything repriced at API list prices and divided by the number of tasks the loop actually closed.

Here's what came out:

ShopClaw spend per 1,000 closed tasks: $3,760 on Opus 5 plus Sonnet 5, $2,189 if the old tokens were repriced at Sonnet 5.5 rates, $515 on Sonnet 5.5 actual. Minus 86 percent.

āˆ’86% per closed task. From $3.76 to $0.51. That's $3,760 → $515 per 1,000 tasks.

And here's the part that surprised me. It's not just the price tag:

  1. Price alone: āˆ’42%. If I'd run the old tokens at Sonnet 5.5 prices, a task would go from $3.76 to $2.19. Not the full 60%, because part of the work was already on Sonnet 5 at the same price.

  2. Volume: another āˆ’76%. The same task then dropped from $2.19 to $0.51, because the new model simply needed far less work to finish it. 5.6Ɨ fewer cache reads. 3.7Ɨ fewer turns per run. And not a single run got cut off by the turn limit (17 out of 103 before, 0 out of 32 after).

Read that twice. Price per token is half the story. Tokens per task is the other half, and a newer model usually needs fewer of them.

How far to trust this (honest version, because I'd rather you check than clap):

  • It's API-equivalent cost, not cash. The loop runs on a Max subscription, so no money actually changed hands. This is what the same work would cost at API list prices (Opus 5 $5/$25, Sonnet 5 and 5.5 $2/$10 per million tokens).

  • The "after" window is short. 3 days, 32 runs, 62 tasks, versus 7.5 days, 103 runs, 152 tasks before. I'll know more in a week.

  • The task mix shifted. After the switch, 56 of 62 closed tasks were small loop fixes (79 of 152 before). Comparing only that same slice gives $3.65 → $0.51 per task, so the drop holds. But that sample is small (16 tasks), so treat it as a sanity check.

  • I didn't measure quality. 152 of 155 tasks closed OK before, 62 of 62 after. That's the log status, not me reviewing the work.

Better and cheaper. In every other industry you pay extra for the new thing.

Next, I’ll show you the tools I use now: where I click in OpenRouter, how I track and A/B test my spending with Cheaper Inference, and the 15-minute check I repeat monthly. It’s the walkthrough I’d want beside me when setting everything up, which is why it’s for paid subscribers.

Tool #1: OpenRouter (the 5-minute tour)

User's avatar

Continue reading this post for free, courtesy of Creators AI.

Or purchase a paid subscription.
Ā© 2026 Creators' AI Ā· Privacy āˆ™ Terms āˆ™ Collection notice
Start your SubstackGet the app
Substack is the home for great culture