Organization

    Your AI costs the most when you prompt poorly

    August 28, 2026·9 min read

    Get control of your AI tokens: why long chats and vague prompts make AI more expensive, and how short answers, focused context and fresh threads save money, electricity and water.

    > 📥 The full guide is available as a PDF (6 pages, Danish). Download it at the top of the page — all advice, examples and sources in one place.

    I hit the Claude Code limit again last week. I choose to see that as a sign of productivity.

    But it made me look at what I was actually paying for. Most of the usage was waste. I was using AI carelessly: long threads, loose prompts, whole files pasted in where one paragraph would have been enough.

    When I wrote about it on LinkedIn, many of you asked for tips. Here they are. These techniques work across the models I use every day: Copilot, Claude, ChatGPT, Mistral and Gemini. Their pricing structures are similar enough that the same habits save money everywhere.

    > Principle. A token is a small piece of a word. You pay per token you send in and per token you get back. It sounds obvious. That is exactly where the whole point is hiding.


    1. What you send is cheap. What you get back is expensive

    With almost every provider, an output token costs more than an input token.

    ModelOutput price compared with input
    Claude5×
    OpenAI6×
    Gemini, depending on model4–6×

    Ratio between output and input prices. Status as of August 2026.

    A model that rambles costs a multiple of one that answers tightly. Ask for the short answer and define the format: "Answer in three bullets", "maximum 200 words", or "give me only the change, not the entire file again". Where possible, also set your chatbot's global preference to short, precise answers.

    Watch the hidden line item: reasoning tokens — the model's "thinking" before the answer — are billed as output. On a heavy analysis, reasoning alone can represent 70–85 percent of the output bill. Turn deep thinking off when the task does not require it.


    2. A long chat is a growing bill

    This is the misconception that costs the most: people think the model "remembers" the conversation for free. It does not.

    Every time you send a message, the model reads the whole conversation again so it can answer. Message 40 also pays for messages 1–39. A long thread therefore becomes more expensive per message as it grows, even if your questions stay short.

    A worked example

    You have been in a chat for an hour. The context now contains 50,000 tokens. You ask a 15-word question. The model rereads all 50,000 tokens to answer. At Claude Opus list price, the input for that one question is roughly $0.25 — and you pay it again with the next question. The short question was never the cost. The baggage was.


    3. The most expensive prompt is the one you send three times

    Because every round trip resends the entire context, the back-and-forth is where the money burns. Five rounds of loose questions cost more than one sharp request. Define the task and deliverable up front:

    1. Role. The role the AI should play.

    2. Minimum context. Only the files and attachments the task needs.

    3. Task. Explain exactly what needs to be done.

    4. Deliverable. The format, audience and constraints.

    Writing that prompt is boring. Not writing it is expensive. You do not need a prompt-engineering course: read those four points three times. That is what you need to know.


    4. Sharp context beats big context

    More context is not better context. Give the model exactly what it needs: the relevant file, paragraph and numbers. Not the whole folder. When people ask why Copilot only accepts 20 files, I smile and explain that they should stay well below 20.


    5. English is the cheapest language for talking to AI

    The major model tokenizers are built around English. English words and endings are often one token; Danish is split into more pieces. The same content therefore takes more tokens in Danish. The difference varies by model, but on some models it can cost up to 40 percent more. Danish letters and long compound words are the worst offenders.

    Danish is still fine for most work. But for a heavy, repeated task where language does not affect the result — such as code, data extraction or structured analysis — try prompting in English. Add "answer in Danish" at the end if the deliverable must be Danish.


    6. Pack up and start again

    This is the habit that has helped me most. When a chat becomes long and sluggish, stop. Ask the AI to create a handover package, then open a fresh chat with it:

    > Create a handover package: a short summary of how far we have come, what has been decided, and what happens next. Write it as a small text file that a new chat can read and continue from.

    You discard the dead weight and keep only what matters. It is cheaper and faster, and answers become sharper because the model does not have to wade through 200 noisy messages to find the thread.


    7. Five more moves that work everywhere

    1. Choose the right model for the task. Use the small, fast model for simple work and the large one for heavy analysis. This alone can cut 40–60 percent from a mixed working day.

    2. Always ask for the short answer. Output is the expensive end. A length or format limit is the easiest saving available.

    3. Turn thinking off when you do not need it. Reasoning tokens are output tokens in disguise.

    4. Start a new chat for a new subject. Do not drag yesterday's conversation into a new task.

    5. Keep stable material at the top. Most providers cache repeated context and offer discounts of up to 90 percent. Put what does not change first in the prompt.


    8. There is another bill besides money

    Tokens mean computation. Computation means electricity and water. Sam Altman has stated that an average ChatGPT query uses about 0.34 watt-hours of electricity and roughly one-fifth of a teaspoon of water. These are OpenAI's own figures from June 2025, not independently verified, so treat them accordingly.

    One question is nothing. Millions of careless questions every day are something else. The good news is that the same habits help both: what lowers your bill also lowers your footprint.

    > Sharp prompting is green prompting.


    9. Three questions to take into Monday

    1. How many of your AI conversations are long threads that should have been packed up an hour ago?

    2. Who has taught your people to ask for the short answer?

    3. What would you save — in money and electricity — if you made it a habit?

    The answer tends to surprise. It is the best opportunity you will get to write a few simple rules for AI use.


    Sources

    Prices and figures were checked against these sources as of August 2026:

    • Claude pricing and prompt caching: platform.claude.com/docs/about-claude/pricing
    • OpenAI pricing: cloudzero.com/blog/openai-pricing
    • Gemini pricing and thinking tokens: cloudzero.com/blog/gemini-pricing
    • Token optimization: tokenoptimize.dev
    • Electricity and water per query: datacenterdynamics.com
    • Language and tokens: arxiv.org/abs/2305.15425 and promptcost.org

    This is general practical guidance, not advice you should use as the sole basis for a decision. All prices are in USD excluding VAT and may have changed; always check the providers' own pricing pages. The electricity and water figures are OpenAI's own, stated by Sam Altman in June 2025, and have not been independently verified. Status as of August 2026.