Turn Optimization into a Habit + a 30-Second Checklist
The four system-layer habits, a 25-item checklist, the tool map and two copy-paste prompts — plus an extension: turn your optimizer into a Skill.
What these notes are
The first two lessons covered how to save within a single round. This one covers how to keep saving — turning the moves into templates, checklists and Skills so you never have to think it through from scratch again.
L5 · System layer: the real savings come from making it a habit
1. Keep measuring: find the biggest leak first
Watch cache hit rates, tool call counts and the token distribution. Acting without data usually means optimizing the layer that did not matter.
2. Template your prompts: sink high-frequency scenarios into files
/prompts
meeting-summary.md
requirement-review.md
data-analysis.md
code-review.md
weekly-report.mdEach time, replace only the variables — {{input}}, {{date_range}}, {{output_format}}.
3. Batch: finish several same-shaped tasks in one go
Analyze the following 5 requirements using the same rules.
Unified output:
ID | Core goal | Risk | Open questions
A. ...
B. ...
C. ...
D. ...
E. ...Handling several same-shaped tasks at once removes repeated background, the fixed prompt and initialization overhead.
4. Stay cache-friendly: stable content first
For systems that support prompt or prefix caching:
Stable content first: system prompt / tool definitions / fixed rules
Variable content after: this question / this datasetDo not change the fixed prefix for no reason. Editing one word can invalidate the cache for the whole prefix.
Optimization is not a one-off diet; it is a measurable loop.
The 30-second self-check: where do you waste tokens
In class, four questions are enough to locate the biggest leak:
| # | Question | Threshold |
|---|---|---|
| 01 | Is the context bloated? | Over 10 rounds / over 15K tokens / several tasks mixed together |
| 02 | Are tools overloaded? | Low cache hits / too many Skills and MCPs / oversized tool definitions |
| 03 | Is the output verbose? | No format specified / lots of explanation / you needed a third of it |
| 04 | Is the model overkill? | Strongest model on simple tasks / redundant prompts |
Find the biggest leak first, then optimize. Whichever of the four sounds most like you is where you start.
The full checklist (work through it after class)
A. Input
- Did I state the goal in one sentence?
- Are the definitions, formulas and boundaries clear?
- Is there a lot of professional, careful, detailed language that changes no behavior?
- Could lists, JSON or tables replace long prose?
- Am I re-pasting information the model already knows?
B. Output
- Is the output format specified?
- Are maximum counts or lengths specified?
- Do I genuinely need background and extra knowledge?
- Did I ask the model not to restate the input?
- Do I stop it when it drifts?
C. Context
- Does this session mix several unrelated tasks?
- Should this long session be summarized into a new one?
- Am I feeding the model an entire PDF, log or codebase?
- Could I search first and read only the relevant part?
- Is a solved debugging process still occupying the context?
D. Models and tools
- Am I using too strong a model or too high a reasoning effort for simple tasks?
- Are all the currently loaded tools, MCPs and Skills necessary?
- Could tools be lazy-loaded?
- Could tool output be filtered by a script first?
- Are several tools duplicating each other?
E. System
- Are high-frequency prompts templated?
- Can same-shaped small tasks be batched?
- Are fixed prompts and tool definitions stable enough for caching?
- Do I regularly look at tokens, latency, cache hits or tool call data?
- Do I record the actual before-and-after difference?
The tool map: which category do you actually need
| What you want to fix | Goal | Example ideas / tools |
|---|---|---|
| Finding code | Fewer file reads, less grep | CodeGraph, Claude Context, Codebase-Memory |
| Compressing output | Swallow fewer diffs, tests and logs | RTK, Headroom |
| Managing context | Automatic trimming, fetch on demand | Context Mode, OpenViking |
| Seeing consumption | Understand budget and cache hits | Tokalator, built-in stats like rtk gain |
| Reducing guesswork and rework | Structured prompts, templates | Prompt libraries |
| Not stuffing whole documents | Search first, then read | RAG / semantic search / on-demand reads |
| Reusing high-frequency capability | Build once, use long term | Skills / prompt libraries |
One line to remember the whole table: find less · read less · store less · see clearly.
Tools change; the ideas do not. Treat those names as examples of what each category looks like, and check the latest official docs before adopting one.
The three closing moves to remember
01 · Feed less
More certain prompts, cleaner context, tools loaded on demand. This step answers how much of the incoming context is actually useless.
02 · Emit less
Cap fields, format and length; if a script can compute it, do not let the model write it out. Output usually costs more than input, so this is the fastest win.
03 · Reuse more
Stable cache prefixes, templated prompts, batched high-frequency tasks. Turn a one-off optimization into something reusable.
Low consumption does not mean lower capability — it means spending tokens where they actually create value.
The memory line
Cut what should be cut, split what should be split, compress what should be compressed.
- Cut: do not generate content that is not needed;
- Split: separate tasks into separate sessions; do not let the context grow forever;
- Compress: filter, summarize and retrieve logs, files and history before handing them to the model.
The final goal is not the lowest possible token count, but: the highest quality result for as few wasted tokens as possible.
Copy-paste prompts
Prompt one · Prompt self-optimization
Optimize the prompt below.
Goal:
Without losing task-critical information or constraints:
1. Delete repeated and useless descriptions;
2. Fill in the key definitions that would otherwise force guessing;
3. Restructure it into a structured format;
4. Specify output format and length;
5. Add no unrelated requirements.
Output:
A. The optimized prompt
B. What was deleted
C. What was added
D. Why this is more efficient
Original prompt:
{{prompt}}Prompt two · Long-conversation compression
See Demo 2 in lesson two — copy it directly.
Extension · Turn a demo into your first Skill
If a demo you ran in class is something you would use every day, make it a Skill instead of re-pasting the prompt each time.
Start with the prompt optimizer — it is the most universal move in the whole course.
The Skill's core instructions
Your task is to optimize the token efficiency of the user's prompt, not to mechanically shorten it.
Principles:
1. Keep the information required to complete the task.
2. Delete repetition, pleasantries and useless role descriptions.
3. Fill in the definitions, boundaries and criteria that would otherwise force guessing.
4. Prefer the goal / input / rules / output / limits structure.
5. Limit unnecessary output.
6. Do not change the user's original goal on your own.
7. Point out gaps when information is insufficient; do not invent business rules.
Output:
## Optimized prompt
...
## Deleted
- ...
## Added
- ...
## Why it is more efficient
- ...Desensitize before publishing
- Remove real client names, company names and internal project names;
- Remove API keys, tokens, accounts and internal URLs;
- Replace sample data with fictional data;
- Check logs, screenshots and config files for private information;
- Make the README state the intended scenarios, inputs, outputs and limits.
For the publishing channel and the detailed flow, see:
- SkillHub: https://www.skillhub.cn/
- SkillHub publishing tutorial: https://www.skillhub.cn/tutorials
Reference links
- SkillHub: https://www.skillhub.cn/
- SkillHub publishing tutorial: https://www.skillhub.cn/tutorials
- RTK (Rust Token Killer): https://github.com/rtk-ai/rtk
Tool versions, install commands, model prices and caching rules all change. Check the latest official documentation for each project or vendor before relying on it.
The 30-second closing
If you remember only one thing today, remember this:
Token optimization is not about counting words — it is about reducing model guesswork,
irrelevant context and useless output.
Make the prompt certain first, then load only the tools you need;
compress long conversations, and turn repeated work into templates and Skills.
Once you start asking "does this information really need to enter the context?",
you have already begun optimizing tokens for real.Pick Tools That Are Just Right + Stop Context Bloat
Model tiering, the four tool principles, the four context-layer moves, 3 live demos and two advanced ideas — the real bulk usually sits here.
Assignment: Find where you waste the most tokens, then fix it
Answer one question — where do you currently waste the most tokens — then install a tool from the course and fix it, with a screenshot.
Tutorials