How to Cut Token Usage in Claude Code and Cursor (2026)
Practical ways to spend fewer tokens with Claude Code and Cursor: measure first, keep context small, match the model to the task, filter what the agent reads, and the open-source tools that help.
Coding agents get expensive in a way that is easy to miss: you pay for what the model reads, and an agent reads a lot. Every request carries the whole conversation so far, plus every file, log and tool result the agent pulled in. A one-line question at the end of a long session can cost more than the first hour of work.
For scale, Anthropic says that across enterprise deployments Claude Code averages about $13 per developer per active day and $150–250 per developer per month, with 90% of users below $30 a day (Claude Code docs, as reported). Cursor says daily Agent users typically spend $60–100 a month and power users $200 or more (Cursor pricing docs). Most of the savings below come from the same few habits in both tools.
1. Measure before you change anything
- Claude Code:
/usageshows the session's token use, cost estimate and prompt-cache hit rate; on Pro, Max, Team and Enterprise plans it also breaks recent usage down by skills, subagents, plugins and MCP servers./contextshows what is filling the context window right now. - Cursor: the usage page in your Cursor dashboard shows spend per model and request.
- Teams on the API: route traffic through a gateway such as LiteLLM to track spend per key, or trace runs with Langfuse. Anthropic's docs mention that several large enterprises use LiteLLM for exactly this.
Look for the obvious culprits first: one session left open all day, a huge instructions file, or tool output (test runs, logs) that dwarfs everything else.
2. Keep the context small
- Clear between unrelated tasks. In Claude Code,
/clearstarts fresh and costs nothing;/renamefirst if you want to/resumelater. In Cursor, start a new chat instead of continuing an old thread. - Compact with intent.
/compact Focus on the failing test and the code changeskeeps what matters. Note that compacting a large session is itself a large request. - Trim your instructions file.
CLAUDE.mdloads into every session. Anthropic suggests keeping it under about 200 lines and moving workflow-specific instructions (PR reviews, migrations) into skills, which load only when used. In Cursor, keep "always apply" rules short and let the rest apply only when relevant. - Hide what the agent should never read. In Cursor, list build output, data dumps and vendored code in
.cursorignore. - Mind idle time. Coming back after the prompt cache has expired reprocesses the whole conversation at full price. The cache lasts an hour on a Claude subscription and five minutes by default on an API key, so a long break is a good moment to
/clearor resume from a summary.
3. Match the model and effort to the task
- Use the mid-size model by default. Anthropic's guidance is that Sonnet handles most coding tasks and costs less than Opus; keep Opus for hard architecture and multi-step reasoning. Switch with
/model. - Lower the effort for simple work. Thinking tokens are billed as output tokens.
/effortlowers the reasoning effort for routine edits. - Give subagents a small model. Set
model: haikuin a subagent's configuration for simple, repetitive jobs. - In Cursor, Cursor's own models come with much more included usage than third-party models, which are billed at their API rates. On Teams and Enterprise, Auto mode has a Cost setting that routes to cheaper models.
4. Shrink what the agent reads
This is where the biggest wins usually are, because tool output is often most of the context.
- Filter command output with a hook. A Claude Code
PreToolUsehook can rewritenpm testorpytestso only failures come back; Anthropic's docs include a ready-made script. The same idea works for logs: return theERRORlines, not the whole file. - Send noisy work to a subagent. Test runs, log digging and documentation lookups can run in a subagent, so only a summary returns to the main conversation.
- Prefer CLI tools to MCP servers where both exist.
gh,awsorgcloudadd nothing to the context until they run. Disable MCP servers you are not using with/mcp. - Compress tool output. Headroom sits between the agent and the model and compresses logs, JSON, search results and files, keeping the originals locally in case the model needs them. Its authors report about 20% fewer tokens for coding agents and far more on repetitive JSON and logs (21–57% in their four published agent scenarios). Treat these as vendor numbers and run its savings report on your own traffic.
- Fetch exact docs instead of whole web pages. Context7 gives the agent version-specific library documentation through MCP, which is both shorter and more accurate than browsing.
- Pack only what matters when you paste code into a chat model. Repomix builds one file from the parts of a repository you choose and shows the token count per file.
| Tool | What it does for your bill | GitHub stars | Price |
|---|---|---|---|
| Headroom | Compresses tool outputs, logs and files before they reach the model, so each run uses fewer tokens. | 74.8K | Free (OSS) |
| Context7 | Gives the agent current, version-specific library docs, so it stops guessing outdated APIs. | 62.8K | Free + $10/mo |
| Repomix | Packs a repository into one AI-friendly file with token counts per file. | 28.8K | Free (OSS) |
| LiteLLM | One gateway for 100+ models with spend tracking, budgets, rate limits and fallbacks. | 60.4K | Free (OSS) |
| Langfuse | Traces every model and tool call with its cost and latency, and runs evaluations. | 35.5K | Free (OSS) |
5. Work habits that avoid wasted runs
- Be specific. "Add input validation to the login function in auth.ts" reads two files; "improve this codebase" reads everything.
- Plan first on big changes. Plan mode (Shift+Tab in Claude Code) agrees on the approach before any expensive edits.
- Stop early. Press Escape as soon as the agent heads the wrong way, and
/rewindto a checkpoint instead of arguing it back. - Give it a way to check itself. Tests or an expected output let the agent catch its own mistakes instead of you paying for another round.
- Be careful with parallel agents. Anthropic notes that agent teams use about 7 times more tokens than a standard session when teammates run in plan mode, because each one has its own context.
Quick checklist
| Habit | Claude Code | Cursor |
|---|---|---|
| See where tokens go | /usage, /context |
Usage dashboard |
| Fresh start between tasks | /clear |
New chat |
| Smaller default model | /model, /effort |
Pick the model; Auto in Cost mode on Teams |
| Keep instructions lean | CLAUDE.md under ~200 lines, skills for workflows |
Short always-apply rules |
| Keep files out of context | Hooks, subagents | .cursorignore |
| Compress tool output | Headroom (headroom wrap claude) |
Headroom (headroom wrap cursor) |
More tools that make agents cheaper or more capable are in agent tooling. For what each coding agent charges in the first place, see AI coding agent pricing.
Agents mentioned in this guide
| # | Agent | Gid Score | Downloads 7d | 7d | 30d | Stars | Latest release | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 3 |
Claude CodeAnthropic |
80 | 14.6M | up 4.4% | down 34.4% | 150K+167/day | 2.1.294yesterday | ||||
| 9 |
CursorAnysphere (owned by SpaceX) |
72 | — | — | — | 33.3K+3/day | — | ||||
| 91 |
|
42 | 143K | up 25.5% | down 20.1% | 74.8K+76/day | 0.40.03d ago | ||||
| 73 |
|
44 | 480K | up 5.6% | down 32.9% | 62.8K+38/day | 0.5.15yesterday | ||||
| 8 |
|
73 | 23.1M | up 1.8% | — | 60.4K+60/day | 1.104.2yesterday | ||||
| 15 |
|
63 | 8.7M | up 11.2% | down 1.9% | 35.5K+39/day | 4.55.0yesterday | ||||
| 66 |
|
46 | 96.9K | down 21.1% | up 14.4% | 28.8K+20/day | 1.18.118d ago | ||||
| No agents match that filter. | |||||||||||