Jul 2026 · The Watch
The scoreboard flipped: intelligence per dollar, not IQ.
An AI lab that isn’t allowed to buy the best chips just made its model serve nearly 7× more people on the same hardware — identical answers, mathematically guaranteed — and then published the recipe for free. Forget the engineering for a second. Look at what it means for your bill.
Why a cost advisor cares about a serving trick
Because every efficiency gain from a lab that publishes its secrets resets the floor price of intelligence for everyone— including the model you’re overpaying for today. This is the clearest signal yet of where the whole market is heading.
Surge pricing was the victory lap, not the cash grab.
DeepSeek announced something no AI lab ever had: time-of-day pricing.Call the API during Beijing business hours, pay double; off-peak, half price. Rush hour for an AI model. It reads like greed — it’s the opposite. You only need traffic management when you’ve built something so fast and cheap the whole world shows up at once. (If you’re a US developer, it’s an accidental gift: Beijing’s peak is the American night — schedule heavy batch jobs in US daytime and you live permanently in the discount window.)
What they actually did (in one breath).
The slow part of AI was never the thinking — it’s the waiting: the chip idles, fetching a mountain of context from slow memory, then does a microsecond of math for one word, then waits again. DeepSeek’s DSpark lets a small, cheap draft model race ahead and guess whole chunks, which the big model then verifies in one pass — final say on every word, so the output is identical, just reached far faster. In their own production: 60–85% faster per user and roughly 7× the throughput on identical hardware, no retraining. Then they MIT-licensed the whole stack — code, checkpoints, training pipeline.
Scruuge’s TLDR
The labs with infinite money mostly kept solving speed the expensive way — stacking more GPUs on the fire. DeepSeek redesigned the fireplace and handed out the blueprints. Constraint wasn’t their handicap; it was their whole strategy.
The honest caveats (because we’re not a hype channel).
- Vendor-reported. The headline 7× / 60–85% numbers come from DeepSeek’s own paper and servers. The open-source benchmarks are reproducible; the production figures aren’t independently confirmed at scale yet. Treat them as vendor-reported until third parties verify.
- It’s data-center medicine, not a magic download. An independent tester ran it at home and got slower results — his draft model was only ~3.5× faster than the target, and this trick generally needs the drafter 10–30× faster before the math pays off. If your prep cook is nearly as slow as the chef, you’ve just hired a middleman.
Why this lands on your bill.
A Deutsche Bank analyst estimated that for most everyday tasks, DeepSeek’s V4 Pro does the job at around 1.5% of the costof top US frontier models — and DSpark cuts its serving cost further. That’s “price gravity,” and it’s real enough that OpenAI is reportedly rethinking its IPO timing. For years the race was scored on one axis: who’s smartest? But intelligence nobody can afford to run at scale is a museum piece. The race that decides which AI ends up inside every app and agent you touch is intelligence per dollar, per second, per watt — and this week a chip-starved lab lapped the field on exactly that scoreboard.
The “best model” is a moving target. Scruuge tracks it.
When the floor price of intelligence resets every few weeks, the model you picked last quarter is probably not the one you should be paying for now. Two ways to keep up: learn what actually separates the models, then run the calculator to see which tier you should be on today.
Sources: DeepSeek’s DSpark paper (MIT-licensed, published on GitHub) and public specs for V4; Deutsche Bank’s June cost estimate; the independent at-home replication as reported. Headline production figures are vendor-reported pending third-party confirmation. Scruuge’s contribution is the cost lens and the honest caveats — not the hype.