I keep hearing people talk about tokens and token spend, blah blah...nodding along like I get it. I didn't. So let's actually start from the beginning: what's a token and why does it cost anything at all?
A token is how language models like Claude or ChatGPT actually read text. Instead of processing a whole text block or images, audio, or video the model breaks it into small chunks — this chopping-up process is called tokenization.
Does every chunk cost the same? Nope, we wish.
Common words are cheap. Words like "the," "and," or "hello" are used so much that they usually get their own single token.
Rare or made-up words are expensive. A word the model hasn't seen much gets split into smaller pieces, so one "word" might cost 3–4 tokens instead of 1. For example: something like "supercalifragilisticexpialidocious" isn't one token, it's several — the model's never seen it as a single chunk so it gets chopped into pieces it has seen before. Try tokencalculator.ai and plug your own name in to see how many pieces it splits into.
Other languages and symbols often cost more. Gen Z might think its cute to use Emojis and unusual punctuation but they frequently take more tokens to say the same thing than plain English does.
Why does this matter for cost?
Every model has a limit on how much text it can hold in a conversation at once — that's called the context window and it's measured in tokens, not words. When you're using an API (rather than a free chat app) you're typically billed per token — both what you send and what the model sends back. More tokens in, more tokens out, more it costs.
So when you hear Anthropic or OpenAI say WOO INTRODUCING the new best model, like GPT-5, it may cost more tokens for the exact same words. Say a prompt works out to 1,000 input tokens and gets a 500-token reply back:
| Model | Cost of that exact same request |
|---|---|
| A smaller, faster model | ~$0.0035 |
| A mid-tier model | ~$0.007 |
| The most capable model | ~$0.0175 |
You're not just paying for text volume, you're paying for how much "thinking power" that model brings — which is exactly why picking a model matters. You don't always need the latest, shiniest release, you just need the most capable one for what you're actually trying to do babes.
Chatting vs. building — not the same bill
So token count and model choice both drive cost — which is exactly why how you prompt matters too. But first, a distinction. If you're just chatting with ChatGPT or Claude through the app, that's not really what this is about. I pay a flat £18/month for Claude and that covers me chatting away within a usage quota — it's not metered per token, so I don't feel any of this day to day.
Earlier tonight I got Claude to add a basic like button to this blog, backed by Supabase. It was not "one prompt and done".
The process:
Scaffold a database table → wire up an API route → wire up the button on the page → test it locally → push it → open a PR → wait for Vercel to build it → discover the env vars weren't set on the deployment yet → add them → discover I'd pasted the publishable key into the service role field (my bad, oops) → fix that → click the button, still nothing → discover the "redeploy" I'd triggered actually rebuilt main, not my preview branch → trigger the right one → wait for it to build → check the logs → finally see a clean 200 → merge.
Every single one of those back-and-forths is tokens. Not just what I typed — what Claude read to figure out what went wrong: files, deployment logs, error traces, all of it. A clean little feature like this, nailed in one go, is maybe a few thousand tokens. A saga like that? A multiple of it. Debugging isn't free, it's just a bill you don't see line-by-line while it's happening.
And that's really why "prompt engineering" is a whole discipline, not just a buzzword:
-
Every word costs money, so vague, rambling prompts waste tokens whether the model needed them or not.
-
Bad prompts cause back-and-forth. An unclear first prompt often means a "no I meant..." follow-up, then another — each round-trip is a fresh set of tokens.
-
How you ask shapes how much you get back. A vague ask like "tell me about marketing" produces a huge, unfocused response. A specific ask tends to produce something shorter, cheaper and more useful. Takes me back to my fav quote by Charlamagne Tha God:
When someone offers to help you, tell them exactly what you want. Don't beat around the bush. If you're not crystal clear about what your ask is, chances are you won't get anything.
Put some thought into those prompts!
- Context adds up fast. Pasting in whole documents or long chat history all counts against the token budget — so careless copy-pasting quietly burns through cost.
The takeaway from all of this: chatting casually barely makes a difference to a quota, but as mentioned above it's a good idea to work on prompting with purpose and not vibes. At scale — a company running thousands of API calls a day, or an automated workflow running on a schedule — it becomes a real cost decision.
And thats why theres all that fuss around token spend!
Till next time xo