ai-token-efficiency-playbook
Health Uyari
- No license — Repository has no license file
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 9 GitHub stars
Code Basarisiz
- rm -rf — Recursive force deletion command in tests/test_check_token_hygiene.sh
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
Drop-in instructions, prompts, and workflows to reduce token usage across Codex, Claude Code, GitHub Copilot, Cursor, and other AI coding agents
AI Token Efficiency Playbook
Drop-in instructions, prompts, examples, and lightweight checks to reduce wasted AI tokens across Codex, Claude Code, GitHub Copilot, Cursor, Gemini, and other AI coding agents.
This project is not just about making AI replies shorter. The main thesis is:
Context is more expensive than reply style.
Most token waste comes from oversized context, repeated memory, noisy logs, and using the wrong model for the job. This playbook focuses on context hygiene and model routing so AI tools stay useful without carrying unnecessary token load.
30-Second Install
Pick your tool, copy the matching file into your repository, and commit it.
| Tool | Copy this file |
|---|---|
| Codex / general coding agents | AGENTS.md |
| Claude Code | CLAUDE.md |
| Gemini CLI / agents | GEMINI.md |
| GitHub Copilot | .github/copilot-instructions.md |
| Cursor | .cursor/rules/token-efficiency.mdc |
cp AGENTS.md /path/to/your-repo/AGENTS.md
For best results, also copy the canonical guidance in guidelines/ and the checklist in checklists/token-efficiency-checklist.md.
What This Provides
Ready-to-copy instruction files and supporting material:
.
├── AGENTS.md
├── CLAUDE.md
├── GEMINI.md
├── .github/
│ └── copilot-instructions.md
├── .cursor/
│ └── rules/
│ └── token-efficiency.mdc
├── prompts/
│ ├── tokensaver-mode.md
│ ├── debugging-mode.md
│ ├── code-review-mode.md
│ └── architecture-mode.md
├── guidelines/
│ ├── token-saving-principles.md
│ ├── coding-agent-guidelines.md
│ ├── context-hygiene.md
│ ├── cli-output-compression.md
│ ├── document-to-markdown.md
│ ├── model-routing.md
│ ├── per-tool-notes.md
│ └── anti-patterns.md
├── examples/
│ ├── before-after-prompts.md
│ ├── bad-vs-good-context.md
│ ├── case-study-ci-log-triage.md
│ └── image-vs-markdown-context.md
├── benchmarks/
│ ├── document-conversion-measurement.md
│ └── dynamic-resource-discovery.md
├── scripts/
│ ├── check-token-hygiene.sh
│ └── estimate-context-size.py
├── tests/
│ ├── test_check_token_hygiene.sh
│ └── test_estimate_context_size.py
├── templates/
│ ├── ci-failure-triage.md
│ ├── document-context-extraction.md
│ ├── handoff-template.md
│ ├── model-routing-decision.md
│ ├── resource-discovery-measurement.csv
│ └── token-savings-measurement.md
└── checklists/
└── token-efficiency-checklist.md
See CHANGELOG.md for the lightweight version history and release notes.
The Standard
The canonical rules live in guidelines/. Tool-specific files such as AGENTS.md, CLAUDE.md, and GEMINI.md are intentionally thin adapters.
AI agents should:
- Use context hygiene before optimizing response style.
- Search before reading large files.
- Read the smallest relevant file range.
- Convert text-heavy documents to searchable Markdown before loading whole documents or screenshot collections.
- Keep the original visual only when layout, charts, diagrams, or image content affect the answer.
- Summarise large logs instead of repeating them.
- Avoid restating unchanged context.
- Keep persistent instruction and memory files short.
- Use the lowest approved model tier capable of the task.
- Escalate early when security, production, architecture, ambiguity, sensitive data, or repeated failure appears.
- Preserve technical accuracy over extreme compression.
For tool-specific context risks, see the per-tool token waste notes.
For common mistakes and fixes, see the token waste anti-patterns catalogue.
For PDFs, Office files, spreadsheets, screenshots, and scans, see the document-to-Markdown context reduction guide.
For lower-model selection and escalation rules, see the practical model-routing guide and the routing decision template.
Why This Matters
Token waste normally comes from four places:
- Oversized input context: whole repos, full files, complete documents, screenshots, and long chat history.
- Noisy tool output: full CI logs, Terraform plans, stack traces, and terminal dumps.
- Repeated instructions: duplicated memory files, repeated project rules, and stale context.
- Poor model routing: using high-cost reasoning models for simple formatting, summarisation, or extraction.
Short replies help, but context hygiene usually saves more.
Practical model routing
Use capability tiers instead of hard-coding model names:
| Tier | Example use |
|---|---|
| Economy / fast | Formatting, extraction, rewriting, selected-text summarisation, deterministic transformations |
| Balanced | Scoped bug fixes, focused code review, ordinary engineering tasks with clear acceptance criteria |
| Advanced reasoning | Security, production, architecture, incidents, ambiguous requirements, broad cross-system work |
Recommended routing flow:
confirm approved data boundary -> classify risk -> choose lowest capable tier -> verify -> escalate when triggered
Do not silently switch provider, tenancy, region, retention policy, approved model family, or data boundary merely to reduce cost. If the agent host cannot change models automatically, it should state the recommended tier rather than pretending a switch happened.
Start with:
Document-to-Markdown workflow
Text-heavy PDFs, Word documents, PowerPoint files, spreadsheets, screenshots, and scans can often be converted to searchable Markdown before AI use. Microsoft MarkItDown is one example of a local conversion utility designed for LLM and text-analysis workflows.
The recommended process is:
convert locally -> inspect -> search -> select relevant sections -> attach a visual only when needed
This avoids providing an entire document when the task needs only a few sections. The original page or image should still be retained when the answer depends on diagrams, charts, colour, spatial layout, UI appearance, handwriting, or complex tables.
Start with:
- Document-to-Markdown guidance
- Image-heavy vs selected-Markdown example
- Document context extraction template
- Document conversion measurement protocol
The playbook makes no fixed token-saving claim for conversion. Measure the complete document, complete Markdown, selected Markdown, and selected-Markdown-plus-visual approaches against the same task.
Emerging benchmark: dynamic resource discovery
Agent clients are beginning to discover skills, tools, agents, and MCP servers on demand instead of preloading every capability into each session. The dynamic resource discovery benchmark defines a vendor-neutral method for comparing those approaches across input context, total tokens, cost, latency, task success, resource selection, and safety.
Use templates/resource-discovery-measurement.csv to capture raw runs. This repository does not claim a saving until comparable, reproducible measurements have been collected.
Example Token-Saver Prompt
Use TokenSaver mode for this task.
Answer first. Read only relevant files. Summarise large outputs. Do not repeat unchanged context. Use the smallest useful verification. Keep final response short unless I ask for detail.
For CI/CD failures, use templates/ci-failure-triage.md to provide the exact failure signal without pasting full raw logs.
For concise task transfers, use templates/handoff-template.md to capture facts, errors, recent changes, tried commands, and open questions without carrying unnecessary context.
Optional Local Check
Run the lightweight hygiene check before committing AI instruction files or logs:
bash scripts/check-token-hygiene.sh
The same check can also run in CI. See docs/token-hygiene-ci.md for the workflow purpose, default thresholds, and tuning guidance.
To produce rough before/after context-size numbers for a case study, use the dependency-free estimator documented in docs/context-size-estimator.md:
python3 scripts/estimate-context-size.py before.md after.md --markdown
Run the helper regression tests before changing file discovery, thresholds, encoding, totals, or output formats:
python3 -m unittest discover -s tests -p 'test_estimate_context_size.py' -v
bash tests/test_check_token_hygiene.sh
The shell suite verifies staged filenames containing spaces, log and context thresholds, instruction-size limits, untracked-file behavior, and the non-Git fallback. GitHub Actions runs both dependency-free suites on pull requests and pushes to main.
With pre-commit:
pip install pre-commit
pre-commit run --all-files
FAQ
Does this make AI less accurate?
No. The goal is to remove low-value context, not useful evidence. Accuracy comes first: keep exact errors, relevant files, commands, assumptions, risks, and verification results.
How do I measure savings?
Use your AI tool's token dashboard where available, or run the same workflow before and after applying the playbook and compare prompt/context size. Record the tool, model, input shape, output shape, and caveats so the result is reproducible. Use templates/token-savings-measurement.md as a starting point.
Is this only for coding agents?
No. It works for AI chat, code review, incident triage, CI debugging, infrastructure-as-code, document review, documentation cleanup, and long-running engineering tasks.
What This Is Not
This is not a benchmark claim that every workflow saves a fixed percentage. Token savings depend on the task, model, tool, document format, conversion quality, and how much context is loaded. The aim is practical reduction without making the AI less useful.
Recommended Use
Use this playbook when working with:
- AI coding agents
- terminal-based AI tools
- large codebases
- text-heavy documents and screenshot collections
- CI/CD logs
- infrastructure-as-code projects
- long debugging sessions
- repeated review or refactor workflows
Contributing
Contributions are welcome. Useful additions include new tool-specific instruction files, before/after examples, workflow patterns, enforcement checks, and measured token-saving case studies.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi