Where Tokens Go + Writing Deterministic Prompts
The token composition of one task, four common kinds of waste, three example prompts and one universal template — correcting the intuition that shorter is better.
What these notes are
This lesson is the foundation. First we balance the books on where the money goes, then we correct one easily flipped intuition: a prompt is not better because it is shorter, it is better because it is more certain.
Everything in the next two lessons — tool choice and context hygiene — builds on the two judgments made here.
In one sentence
Token optimization is not about writing a few words less — it is about making every token that enters the model context useful.
One metric captures it:
Token efficiency = useful tokens ÷ total tokens spentThe closer to 1, the more efficient.
What you are actually trying to reduce is these seven things: model guesswork, irrelevant history, repeated background, excessive tool definitions, verbose logs, useless output, and rework.
Where tokens actually go
In a single task, the context usually carries all of this at once:
System prompt + user input + conversation history + file / retrieval results
+ tool / MCP / skill definitions + tool return values + model outputAudit it as six blocks:
| Where it goes | What it looks like | What runs away first |
|---|---|---|
| System prompt | Role description, fixed rules | Gets longer and longer, all restating the same thing |
| User input | This task and its materials | Re-pasting information the model already knows |
| Conversation history | Previous rounds | One session mixing several unrelated tasks |
| Files / retrieval | Whole documents dropped in | Handing over an entire PDF |
| Tool definitions | Skill and MCP schemas | The more you install, the higher the hidden cost |
| Tool return values | Command output, logs | Thousands of lines, a few dozen that matter |
Plus a seventh block: model output — you asked for a conclusion and received a whole essay of background.
Two things worth remembering first
First, output tokens are usually more expensive than input tokens. So limiting unnecessary output is the fastest win available.
Second, tool schemas also occupy the context. A pile of Skills and MCPs you never call still costs you before any work starts — more tools, more hidden cost.
The price ratio between input and output tokens varies by model, vendor and time — check current pricing. The stable, effective approach is to eliminate inputs and outputs that carry no business value.
Four common kinds of waste
| Waste | What it looks like | What to do |
|---|---|---|
| A long prompt with an unclear goal | Plenty of background, no definitions, boundaries or output requirements | Fill in definitions and decision criteria, cut the pleasantries |
| A context that keeps growing | One session with a weekly report, a spreadsheet, code and a marketing plan | One task one session, or write a handoff summary |
| Too much tool output | Thousands of log lines, a few dozen that matter | Filter with a script first, then hand it to the model |
| Output beyond what is needed | You asked for a conclusion plus three reasons and got background plus tangents | Specify fields, format and maximum counts |
See which ones you recognize — the second lesson gives each one a matching action.
L2 · Input layer: not shorter, but more certain
Four actions
| # | Action | What it actually means |
|---|---|---|
| 1 | State the goal | Give the goal, definition, formula, boundaries and decision criteria directly |
| 2 | Structure the input | Prefer goal / input / rules / output / limits over long prose |
| 3 | Slim the system prompt | Delete repeated role descriptions, synonyms and weak constraints |
| 4 | Limit the output | Specify fields, format, maximum counts, and whether explanation is needed |
The point: shrink the space where the model has to guess, and shrink rework. Rework is the most expensive kind of waste — it turns every earlier token into a sunk cost.
Three example prompts
Example A · From a guessed metric to an executable formula
❌ Vague:
Help me calculate the retention rate for each group.The model can only guess: day 1, day 7 or day 30? What is the denominator? How is active defined?
✅ Certain:
Task: calculate the 7-day retention rate for each group.
Formula:
7-day retention = users still active on D+7 / new users on D × 100%
Definitions:
- Active: at least one app_open that day
- New: first occurrence of register
- Grouping: group_id
- Date range: 2025-06-01 to 2025-06-30
Output:
group_id | new users | D+7 active users | 7-day retention
Output the result table only; mark missing data as [missing], do not fill in values.Notice one thing: the optimized version has more words, but there is nothing left for the model to guess. That is why certainty is worth more than shortness.
Example B · Slimming a system prompt
❌ Before:
You are a very professional, very senior, highly experienced data analysis expert.
Always stay professional, accurate and well organized.
Answers must be concise, do not ramble, do not add unnecessary explanations.
Do not repeat information the user already knows, and make sure to use Chinese.
Do not make things up when information is uncertain.✅ After:
You are a data analysis assistant.
Rules:
1. Output only the content and format the user asked for.
2. Mark missing or unverifiable information as [unconfirmed]; do not invent it.
3. Use Chinese.Principle: delete what looks professional, keep what actually changes behavior.
Words like highly senior, professional or well organized do not change a single output; output only the requested format and mark missing data as [unconfirmed] do.
Example C · Limiting the output
Analyze the requirement below and output only:
1. Core goal: one sentence
2. Users: at most 3 categories
3. Must-have features: at most 5 items
4. Biggest risks: at most 3 items
5. Open questions: at most 3 items
Do not restate the original requirement, do not write an introduction.Writing at most N items is the most direct way to cap output.
The universal token-saving prompt template
Task:
[one sentence on what must be done]
Input:
[only the information required to do it]
Rules:
1. [key decision rule]
2. [key boundary]
3. When information is insufficient, mark it — do not assume.
Output:
[fields / JSON / table / Markdown]
Limits:
- at most [N] items
- do not restate the input
- no introduction or background
- no unrelated extrasMemory line: goal / input / rules / output / limits.
Demo 1 · The prompt diet: reduce guessing, do not mechanically delete words
Original prompt
Help me analyze this requirement, see what problems it has, and give some suggestions.Optimized prompt
You are a product requirement review assistant.
Task:
Review whether the requirement below is ready to enter development.
Check only:
1. Is the goal clear
2. Are the users clear
3. Are the acceptance criteria complete
4. Are there critical dependencies
5. Is there obvious ambiguity
Output:
- Verdict: ready for development / needs more detail
- Missing information: at most 5 items
- Risks: at most 3 items
- Questions the product owner must answer: at most 5 items
Do not restate the requirement.
Do not suggest new features beyond the requirement.
Requirement:
{{paste the requirement}}Four things to watch live
- Did it restate the original text
- Did it stay generic and vague
- Did one round produce something usable
- Did it reduce the number of follow-up prompts
Demo takeaway: short does not equal efficient; high certainty does.
Lesson recap
| Remember | In one sentence |
|---|---|
| What optimization means | Make every token in the context useful, not write fewer words |
| Where to start | Limit unnecessary output first, then make the prompt certain |
| How to be certain | Goal / input / rules / output / limits |
| What to delete | Words that look professional but change nothing, plus re-pasted known information |
Next lesson moves up two layers: choosing the right model and tools, and stopping the context from growing without end.
Save & Use Tokens Efficiently
A five-layer token optimization model with 3 live demos — it is not about writing fewer words, it is about making every token in the context useful.
Pick Tools That Are Just Right + Stop Context Bloat
Model tiering, the four tool principles, the four context-layer moves, 3 live demos and two advanced ideas — the real bulk usually sits here.
Tutorials