Skip to content
Agents tracked: 283 Downloads (7d): 261M up 5.8% GitHub stars: 6.2M VS Code installs: 151M Releases (7d): 289 Agent status: 1 with issues Updated Oct 9, 2026

How to Cut Token Usage in Claude Code and Cursor (2026)

Practical ways to spend fewer tokens with Claude Code and Cursor: measure first, keep context small, match the model to the task, filter what the agent reads, and the open-source tools that help.

Coding agents get expensive in a way that is easy to miss: you pay for what the model reads, and an agent reads a lot. Every request carries the whole conversation so far, plus every file, log and tool result the agent pulled in. A one-line question at the end of a long session can cost more than the first hour of work.

For scale, Anthropic says that across enterprise deployments Claude Code averages about $13 per developer per active day and $150–250 per developer per month, with 90% of users below $30 a day (Claude Code docs, as reported). Cursor says daily Agent users typically spend $60–100 a month and power users $200 or more (Cursor pricing docs). Most of the savings below come from the same few habits in both tools.

1. Measure before you change anything

  • Claude Code: /usage shows the session's token use, cost estimate and prompt-cache hit rate; on Pro, Max, Team and Enterprise plans it also breaks recent usage down by skills, subagents, plugins and MCP servers. /context shows what is filling the context window right now.
  • Cursor: the usage page in your Cursor dashboard shows spend per model and request.
  • Teams on the API: route traffic through a gateway such as LiteLLM to track spend per key, or trace runs with Langfuse. Anthropic's docs mention that several large enterprises use LiteLLM for exactly this.

Look for the obvious culprits first: one session left open all day, a huge instructions file, or tool output (test runs, logs) that dwarfs everything else.

2. Keep the context small

  • Clear between unrelated tasks. In Claude Code, /clear starts fresh and costs nothing; /rename first if you want to /resume later. In Cursor, start a new chat instead of continuing an old thread.
  • Compact with intent. /compact Focus on the failing test and the code changes keeps what matters. Note that compacting a large session is itself a large request.
  • Trim your instructions file. CLAUDE.md loads into every session. Anthropic suggests keeping it under about 200 lines and moving workflow-specific instructions (PR reviews, migrations) into skills, which load only when used. In Cursor, keep "always apply" rules short and let the rest apply only when relevant.
  • Hide what the agent should never read. In Cursor, list build output, data dumps and vendored code in .cursorignore.
  • Mind idle time. Coming back after the prompt cache has expired reprocesses the whole conversation at full price. The cache lasts an hour on a Claude subscription and five minutes by default on an API key, so a long break is a good moment to /clear or resume from a summary.

3. Match the model and effort to the task

  • Use the mid-size model by default. Anthropic's guidance is that Sonnet handles most coding tasks and costs less than Opus; keep Opus for hard architecture and multi-step reasoning. Switch with /model.
  • Lower the effort for simple work. Thinking tokens are billed as output tokens. /effort lowers the reasoning effort for routine edits.
  • Give subagents a small model. Set model: haiku in a subagent's configuration for simple, repetitive jobs.
  • In Cursor, Cursor's own models come with much more included usage than third-party models, which are billed at their API rates. On Teams and Enterprise, Auto mode has a Cost setting that routes to cheaper models.

4. Shrink what the agent reads

This is where the biggest wins usually are, because tool output is often most of the context.

  • Filter command output with a hook. A Claude Code PreToolUse hook can rewrite npm test or pytest so only failures come back; Anthropic's docs include a ready-made script. The same idea works for logs: return the ERROR lines, not the whole file.
  • Send noisy work to a subagent. Test runs, log digging and documentation lookups can run in a subagent, so only a summary returns to the main conversation.
  • Prefer CLI tools to MCP servers where both exist. gh, aws or gcloud add nothing to the context until they run. Disable MCP servers you are not using with /mcp.
  • Compress tool output. Headroom sits between the agent and the model and compresses logs, JSON, search results and files, keeping the originals locally in case the model needs them. Its authors report about 20% fewer tokens for coding agents and far more on repetitive JSON and logs (21–57% in their four published agent scenarios). Treat these as vendor numbers and run its savings report on your own traffic.
  • Fetch exact docs instead of whole web pages. Context7 gives the agent version-specific library documentation through MCP, which is both shorter and more accurate than browsing.
  • Pack only what matters when you paste code into a chat model. Repomix builds one file from the parts of a repository you choose and shows the token count per file.
Tool What it does for your bill GitHub stars Price
Headroom Compresses tool outputs, logs and files before they reach the model, so each run uses fewer tokens. 74.8K Free (OSS)
Context7 Gives the agent current, version-specific library docs, so it stops guessing outdated APIs. 62.8K Free + $10/mo
Repomix Packs a repository into one AI-friendly file with token counts per file. 28.8K Free (OSS)
LiteLLM One gateway for 100+ models with spend tracking, budgets, rate limits and fallbacks. 60.4K Free (OSS)
Langfuse Traces every model and tool call with its cost and latency, and runs evaluations. 35.5K Free (OSS)

5. Work habits that avoid wasted runs

  • Be specific. "Add input validation to the login function in auth.ts" reads two files; "improve this codebase" reads everything.
  • Plan first on big changes. Plan mode (Shift+Tab in Claude Code) agrees on the approach before any expensive edits.
  • Stop early. Press Escape as soon as the agent heads the wrong way, and /rewind to a checkpoint instead of arguing it back.
  • Give it a way to check itself. Tests or an expected output let the agent catch its own mistakes instead of you paying for another round.
  • Be careful with parallel agents. Anthropic notes that agent teams use about 7 times more tokens than a standard session when teammates run in plan mode, because each one has its own context.

Quick checklist

Habit Claude Code Cursor
See where tokens go /usage, /context Usage dashboard
Fresh start between tasks /clear New chat
Smaller default model /model, /effort Pick the model; Auto in Cost mode on Teams
Keep instructions lean CLAUDE.md under ~200 lines, skills for workflows Short always-apply rules
Keep files out of context Hooks, subagents .cursorignore
Compress tool output Headroom (headroom wrap claude) Headroom (headroom wrap cursor)

More tools that make agents cheaper or more capable are in agent tooling. For what each coding agent charges in the first place, see AI coding agent pricing.

Agents mentioned in this guide

# Agent Gid Score Downloads 7d 7d 30d Stars Latest release
3
Claude CodeAnthropic
80 14.6M up 4.4% down 34.4% 150K+167/day 2.1.294yesterday
9
CursorAnysphere (owned by SpaceX)
72 — — — 33.3K+3/day —
91
Headroom NewHeadroom Labs
42 143K up 25.5% down 20.1% 74.8K+76/day 0.40.03d ago
73
Context7 NewUpstash
44 480K up 5.6% down 32.9% 62.8K+38/day 0.5.15yesterday
8
LiteLLM NewBerriAI
73 23.1M up 1.8% — 60.4K+60/day 1.104.2yesterday
15
Langfuse NewLangfuse
63 8.7M up 11.2% down 1.9% 35.5K+39/day 4.55.0yesterday
66
Repomix NewRepomix
46 96.9K down 21.1% up 14.4% 28.8K+20/day 1.18.118d ago