Save & Use Tokens Efficiently
A five-layer token optimization model with 3 live demos — it is not about writing fewer words, it is about making every token in the context useful.
Token optimization is not about saying a few words less — it is about making every token that enters the model context useful. This course walks the five-layer model (cognition → input → tools → context → system) starting from whichever layer wastes the most, with 3 demos run live.
What you take home
| 1 metric | 1 structure | 5 layers | 3 demos |
|---|---|---|---|
| Useful tokens ÷ total tokens | Goal / Input / Rules / Output / Limits | From knowing where the money goes to making it sustainable | Prompt diet, context slimming, log compression |
Every session delivers three things: hands-on practice, slides and notes, and homework.
What you have probably run into
- A single conversation gets more expensive the longer it goes, and the AI starts drifting;
- You write a long prompt and still get something long and empty, requiring several more rounds;
- Even simple tasks run on the strongest model, and the bill climbs for no clear reason;
- You paste thousands of log lines and only a few dozen matter.
The problem is usually not that you typed too much — it is that the context holds too much that has nothing to do with this task.
What you will learn
Where tokens go + how to write deterministic prompts
First, work out which blocks a single task spends tokens on, and recognize the four most common kinds of waste. Then move into the four input-layer actions, using three examples (turning a guessed metric into an executable formula, slimming a system prompt, limiting output) and one universal template to replace shorter is better with more certain is better.
Pick tools that are just right + stop context bloat
Model tiering and script-first filtering; the four tool principles and RTK; the four context-layer moves (one task one session, handoff summary, fetch files on demand, let scripts do deterministic filtering); three live demos covering prompt dieting, long-conversation compression, and tool-output compression; plus two advanced ideas: Context Mode and CodeGraph.
Turn optimization into a habit + a 30-second checklist
The four system-layer habits (keep measuring, template it, batch it, stay cache-friendly); a 25-item self-check list and a four-question quick pass; the tool map and two copy-paste prompts; three closing moves (feed less, emit less, reuse more); and an extension: turn your prompt optimizer into your first Skill.
Who this is for
Course structure
This course has 3 lessons and 1 assignment, climbing the five-layer token optimization model from the bottom up. No need to fix everything at once — start where the waste is worst.
Where tokens go + how to write deterministic prompts
The token composition of a single task and the two things worth remembering first; four common kinds of waste; the four input-layer actions; three example prompts and one universal template; Demo 1: the prompt diet.
Pick tools that are just right + stop context bloat
Model tiering and script-first filtering; the four tool principles and RTK; the four context-layer moves; Demo 2 and Demo 3; advanced ideas — keep raw data out of the context with Context Mode, and stop Agents from digging through the repo with CodeGraph.
Turn optimization into a habit + a 30-second checklist
The four system-layer habits; the 25-item checklist and four-question quick pass; the tool map; three closing moves and the memory line; two copy-paste prompts; and an extension: turn a demo into your own Skill and publish it.
Core principle
In one sentence: get the highest quality result with as few wasted tokens as possible.
Optimization is not a one-off diet but a measurable loop: find the biggest leak first, then fix it, then measure again.
What you are really reducing: model guesswork, irrelevant history, repeated background, excessive tool definitions, verbose logs, useless output, and rework.
Remember the line: cut what should be cut, split what should be split, compress what should be compressed.
Assignment: Run One Real Task Through the Formula
One sentence stating the result, three rounds of dialogue records, one sentence of retrospective — all three required; missing one means it does not count.
Where Tokens Go + Writing Deterministic Prompts
The token composition of one task, four common kinds of waste, three example prompts and one universal template — correcting the intuition that shorter is better.
Tutorials