← AI Cost Watch

Jul 2026 · The Watch · Cost Engineering

Cheap execution is the floor, not the strategy.

We just showed you how to cut a feature’s bill ~68% by routing to cheaper models. Here’s the honest second half that nobody else selling you a cost tool will say: that’s the floor.

Why a cost advisor is telling you cost isn’t the whole game

Because we profit from your clarity, not your confusion. Selling you endless routing tweaks while the real leverage sits one layer up would be exactly the kind of thing Scruuge exists to call out. Cut the waste — absolutely. Then read this.

Better tools, the same results.

The tools keep getting better, prices keep dropping, everyone can produce more than ever — and somehow the output all starts to look the same. Yours, your competitor’s, half your feed.

That’s not a tooling problem. AI made executioncheap, and when execution gets cheap the value doesn’t disappear — it moves up a layer, from doing the task to knowing which task is worth doing. Here’s exactly where it moved, in one true story.

The $40 test.

Mitchell Hashimoto — co-founder of HashiCorp, creator of the Ghostty terminal, about as good an engineer as exists — spent days testing the new frontier model against cheap ones on ordinary work: “implement this feature,” “build this thing.” The stuff on everyone’s list. All three models produced equally acceptable output. The budget model cost under a dollar; the frontier model, ~$9. Same work, same quality, 9× the price.

Stop there and the frontier model looks like a rip-off — which is exactly the hot take flooding every feed right now: route everything to cheap models. And you should route a lot of execution to cheap models. But Hashimoto ran one more test.

He handed the frontier model a problem the cheap ones couldn’t touch: optimizing a gnarly piece of systems code he’d written himself. Two hours. $40. A level of performance he says he couldn’t have reached on his own.

Scruuge’s TLDR

Who assigned that $40 task? No backlog. No sprint. No PM. It didn’t exist as a task until an expert suspected something new was possible and spent money to find out. AI can only do work someone imagined — the ceiling was never the model or the price. It’s the size of your list of things you know how to ask for.

The floor and the ceiling.

Look at the first half again: cheap model tiesfrontier model. That’s not a fact about the models — it’s a fact about the task.“Implement this feature” is work everyone already knows how to ask for, and the work everyone knows how to ask for is exactly where the models — and the people using them — have all converged. That’s your sameness.

So there are two layers, and you need both:

  • The floor — cheap execution. Routing (our 68% piece) and compiling repeats (compiled AI). Real savings — and about to be table stakes. Every competitor with the same insight gets them too.
  • The ceiling — frontier imagination. The surgical use of the best model on the questions that change what the execution layer is even building. The $40 questions no one else thought to ask.

Cheap execution is a great engine. Frontier imagination is where you steer. They’re not in competition — imagination is the multiplier on execution. And the cheaper execution gets, the more it commoditizes, the more valuable every frontier-posed question becomes.

Imagination isn’t artistry — it’s touch.

The word oversells it. This isn’t a gift artists have and analysts don’t. Hashimoto could pose that question because he has thousands of hours inside these models — he knows where the capability line moved, not from a benchmark chart but from instinct, from touch. You can’t imagine with capabilities you haven’t handled.Nobody invents a use for a tool they’ve only read a summary of.

And here’s where most of us quietly sabotage ourselves: we come to AI with cost savings in the back of our head and a fixed task list in front of us, asking “can this do my existing work faster or cheaper?” That’s pointing the telescope at the ground. The questions that find new territory sound different: what can this do that I could never even ask for before?

The one-question audit

Has your task list actually changed in the last 12 months? The last 6? The last 3?

Has what you ask AI to doshifted — or are you running the same old list faster and cheaper and calling it transformation? If it’s the old list, nothing is wrong with your tools. You have an imagination shortage, and you’re about to spend a lot optimizing execution in a commoditized market while your differentiation quietly evaporates.

Redesign the building, don’t bolt on a motor.

When factories electrified, the tech worked on day one — but the productivity payoff took decades, because managers kept the steam-era layout and just bolted an electric motor where the steam engine used to be. Same building, new power source, barely any gain. The payoff came when a new generation redesigned the building around what cheap, distributed motors made possible.

AI is the same kind of technology, and most companies are making the same mistake: bolting it onto the old layout — running the existing task list through cheaper models and reporting the savings. Real savings. Also available to every competitor. Table stakes.

Here’s what redesigning the building looks like when it works. Stripe ran a migration across 50 million lines of code in a single day— work estimated at two-plus months. The impressive number isn’t “a day.” It’s the yearsStripe spent first, building the test coverage and review systems that could verify that many changes safely. Point the same model at a company that hadn’t done that work and you don’t get a one-day migration — you get 50 million lines of changes nobody can approve. They built the building, then harvested it with a frontier model.

You can’t hire your way out of it.

The tempting shortcut is to hire one imaginative “AI visionary.” But that $40 job needed imagination, deep context, and permission to ask — all in one head. A new hire brings the imagination and none of your context, and imagination only fires next to context. Your context is spread across everyone who actually does the work.

So the job isn’t hiring imagination — it’s manufacturing it: putting the people who hold context in contact with capable models, and giving them permission to make bets. The org version of the test: who on your team is allowed to pose a $40 question to a model today without asking anyone?If it’s nobody, that’s an imagination constraint — and it was never about the price of the model.

Cut the floor with us — then go find your ceiling.

Scruuge’s job is the floor: the calculator shows what you’re overpaying in two minutes, and the $999 Assessment maps bothlayers — where you’re wasting spend, and where you’re optimizing a commoditized task list instead of pointing the frontier at a question that would 10× you. We’ll cut your bill. We’ll also tell you, honestly, when cost was never your real constraint.

Built on a widely-shared essay on where AI value moves. The $40 test is Mitchell Hashimoto’s publicly-shared experiment; the 50-million-line/one-day migration is per Stripe’s public account; the factory-electrification parallel is a well-documented economic history. Scruuge’s contribution is the floor-and-ceiling frame and the honest admission that cost-cutting is only the floor.