Pick Tools That Are Just Right + Stop Context Bloat
Model tiering, the four tool principles, the four context-layer moves, 3 live demos and two advanced ideas — the real bulk usually sits here.
What these notes are
The previous lesson covered how to write. This one moves up two layers: what you pick to do the work (L3, tools) and how to stop the context from growing without end (L4, context).
A quick spoiler: the real bulk sits in the context layer.
L3 · Model and tool layer: not the strongest, but just right
Model tiering: read the task first, then pick the tier
| Task type | What to use |
|---|---|
| Classification, extraction, format conversion, simple summaries | Fast / lightweight model |
| Ordinary writing, analysis, office work | Mid-tier model |
| Complex code, hard research, complex reasoning | Strong model / higher reasoning effort |
| Not sure | Try the cheaper option first, escalate on failure |
That last row is not about agonizing — start cheap by default and escalate only if it fails. It is more economical than opening with the strongest configuration, and it rarely changes the quality of the result.
Script first: if a machine can compute it deterministically, do not let the model read and then compute
Let machines do deterministic filtering first; hand only the parts that need judgment to the model.
Lookups, filtering, calculations and caching all belong here. Asking a model to count a column or total a table is slow, expensive and error-prone; a script does it in one line.
The four tool principles
Minimal install · enable on demand · lazy load · exit unrelated tools when the task endsTools, MCPs and Skills put names, descriptions and parameter schemas into the context. The more tools and the more complex their definitions, the higher the potential cost — even if you never call them.
RTK: filter before CLI output enters the context
RTK (Rust Token Killer) filters and compresses command-line output before it enters the LLM context.
Check whether it is installed first:
rtk --version
rtk gainOnce installed, rtk gain shows the savings statistics. The project publishes benchmarks showing large reductions for common development commands, but real results vary by project, command and workflow.
The point of the class is not to memorize a percentage, but to understand: do not hand the model an entire raw machine output.
Raw data → machine filter / compress → key results → LLM judgmentDemo 3 · Tool output compression: teach the Agent to eat less log
A. Manual filtering (the simplest option — learn this one first)
npm test 2>&1 | tail -n 100
grep -iE "error|failed|exception" app.log | tail -n 100B. RTK (optional, for AI coding scenarios)
See the previous section. Once installed, use rtk gain to compare the savings.
What to watch: compare the length before and after filtering, then check whether the model's judgment got worse.
L4 · Context layer: the real bulk usually sits here
8.1 One task, one session
One goal → one session → done → start a new session for the next taskDo not keep mixing unrelated tasks into the same context. Code, deployment, weekly reports and marketing plans in one session means every round pays for all the history before it — and the AI gets pulled off course by stale topics.
8.2 Compress long sessions into a handoff summary
Compress the current session into the minimum context needed to keep working.
Keep only:
1. Current goal
2. Confirmed facts
3. Completed items
4. Current approach / key decisions
5. Open issues
6. Next step
7. File names, variable names, interface names and commands that must be preserved
Delete:
- Pleasantries
- Rejected options
- Repeated explanations
- Intermediate trial and error
- Anything unrelated to what comes next
Requirement: a new session must be able to continue from this summary alone.8.3 Fetch files on demand, do not stuff in the whole book
Search first → find the relevant section or page → read only that → answerThe core rule: take what you need, not everything you have.
8.4 Let scripts handle deterministic filtering
See Demo 3. Never let raw logs appear in the model context.
Demo 2 · Context slimming: from a long conversation to a minimal handoff summary
Demo 1 was a diet for a prompt; this one is a diet for a conversation. Use this extended prompt to make the old session produce its own handoff summary:
Compress the current project into the minimum context needed to keep working in a new session.
Must keep:
- Goal
- Confirmed requirements
- Current architecture / approach
- Completed work
- Open issues
- Key names of files, interfaces, variables
- Next step
Must delete:
- Pleasantries
- Repeated explanations
- Rejected options
- Full debugging process for already-solved errors
- History unrelated to the next step
Finally add:
[Suggested next prompt]
so I can copy it into the new session and continue directly.Steps
Paste the prompt above into the old session and let the model produce the handoff summary.
Paste only the summary, the current task, and the files genuinely needed, then continue working.
Three things to watch: whether the input got noticeably shorter; whether the AI still knows the goal, the key decisions and the next step; whether the noise from old topics is gone.
Advanced idea 1 · Keep raw data out of the context
One family of tools takes the approach that raw data should never enter the context at all — only a distilled result comes back. Typical benefits:
| Capability | Order of magnitude | What it does |
|---|---|---|
| Context saving | 98% | Raw data stays in a sandbox; only distilled results return |
| Session continuity | 6× | Tracks file edits, tasks and errors so a compressed session can resume |
| Think in code | 100× | Lets the model generate a calculation script instead of reading the whole dataset |
The scenarios that fit all share one trait: large input, small conclusion — Playwright page captures, CSV statistics, log triage, long-document extraction.
Specific tool names and multipliers change between versions. What matters is the idea: compute in the sandbox; put only the conclusion into the context.
Advanced idea 2 · Stop the Agent from digging through the repo every time
This is the most typical waste in AI coding scenarios.
| Traditional exploration | With a code graph | |
|---|---|---|
| How | grep → read → ls → grep again | Build a local code graph, then query by structure |
| Result | The Agent pieces structure together from file and function names | Call relationships and impact surface are queried directly |
| Cost | A large share of tokens spent finding code | Tokens spent changing code |
Common capabilities: search · callers · callees · impact · context · affected.
When it is worth the most
- Fixing bugs in an old project: check callers, callees and impact before touching anything, so you stop acting on a local view alone;
- Asking questions about a large repo: questions like how does the order flow work or where does the cache expire are exactly what structural exploration answers;
- CI test selection: use affected as the first filter, which suits repos with many tests and slow full runs.
One boundary to respect: static analysis is not runtime truth. Dynamic imports and reflection create blind spots, so do not treat graph results as the only evidence.
Lesson recap
| Remember | In one sentence |
|---|---|
| Model choice | Start cheap, escalate on failure; do not open the strongest for simple tasks |
| Tools | Minimal install, enable on demand, lazy load, exit when the task ends |
| Logs and command output | Filter with a script before it enters the context |
| Long sessions | Summarize and start fresh; do not let the context grow without end |
| Files | Search first, then read — take what you need |
| Advanced | Compute in the sandbox; use a code graph instead of digging through the repo |
The next lesson turns these moves into daily habits and hands you a 30-second checklist.
Where Tokens Go + Writing Deterministic Prompts
The token composition of one task, four common kinds of waste, three example prompts and one universal template — correcting the intuition that shorter is better.
Turn Optimization into a Habit + a 30-Second Checklist
The four system-layer habits, a 25-item checklist, the tool map and two copy-paste prompts — plus an extension: turn your optimizer into a Skill.
Tutorials