Improved Token Efficiency for Longer Agent Runs
Anysphere, Wednesday, September 23rd, 2026
Cursor cut token costs 7% on long agent runs via prompt trimming, dynamic tool loading, and caching.
Cursor's engineering team describes a set of optimizations that reduced token costs for longer agent runs by 7% without sacrificing output quality.
Changes include trimming the system prompt by 66%, dynamically loading tools only when needed to cut static context by 60%, and using explicit cache breakpoints to reduce cold cache misses by 20%. Additional savings came from compressing file line numbers and using subagents to isolate context between agent instances.
The result is that token consumption scales more slowly than the complexity of the tasks being completed.