AI token limits are quietly shrinking while prices stay the same, and how you prompt is starting to matter for your wallet.
A few days ago, I was sitting at my desk, wanted to quickly have Claude solve a somewhat complex task, and—bam.*“Daily limit reached.”*There it was again, that invisible ceiling.

We’ve all been extremely spoiled over the past few months. The big providers have burned through billions to feed us fascinating tools. But even AI companies have to turn a profit at some point. And since they know there would be an outcry online if they suddenly doubled subscription prices from $20 to $40, they’re employing a tactic straight from the supermarket shelf.
Instead of raising the price, they’re secretly shrinking the contents. Shrinkflation.
Only here, it’s not the bag of chips that’s shrinking, but the currency of AI: the tokens. And it’s happening completely under the radar.
I see this happening everywhere. A great custom agent in Notion that cleans up in the background every day? Great—until the free tokens suddenly run out by mid-month. Even with Adobe, I’m increasingly hitting the limit of my available credits. And I ask myself quite pragmatically: How long will tools like Google NotebookLM actually remain this generous and free?

In Notion, the token credits are used up pretty quickly.
It’s not just the providers’ tactics that are interesting, but also our own behavior. We’ve gotten into the habit of blindly dumping huge, completely unstructured PDFs into the chat. “Make a summary!”
That’s convenient, but it’s a massive token waster. The AI has to process half the layout, headers, and endless amounts of unimportant text. That costs computing power—and thus our tokens.
The solution is actually simple: we need to learn to prompt more efficiently again.
That means, for example, that we start using structured text formats like Markdown instead of uploading PDFs. This saves a huge number of tokens, and the AI understands the text better anyway.
And above all, we need to ask ourselves: Does it really always have to be the biggest and smartest algorithm? For simple tasks, such as proofreading a text or writing a short summary, an older, much faster, and token-friendly model is often perfectly sufficient. Not every nail needs a sledgehammer.
For those who no longer want to be held back by all these token budgets and limits, there’s an exciting alternative: escaping the cloud.
With tools likeOllama, you can bring AI directly to your own computer. For many everyday tasks, a local LLM (Large Language Model) is absolutely sufficient. The nice side effect? You are your own token limit. And the data never leaves your desk.
More efficient prompting will become a basic skill if we don’t want to constantly find ourselves standing in front of closed digital doors. And ultimately, it’s also about our wallets.
How about you? Which tool did you hit the limit with most recently? And have you ever tried running a local AI like Ollama on your own machine?