How to Save on AI Tool Spend & Tokens — Three Levers That Compress Unoptimized Cost to 20-30%
AI bills balloon because output tokens cost 5-6x more than input, context is resent in full every turn, and sub-agents fire multiple times in the background. This article shows how to combine "three levers" — prompt caching (-60 to 90%), model selection (-50 to 80%), and output budget (-30 to 60%) — to compress unoptimized cost to 20-30%, drawing on Anthropic's official guidance, industry research, and real operational data. Covers the early-2026 cache TTL shortening (60 min → 5 min) trap, context management with /compact, the multi-agent 15x token trap, monitoring and billing alerts, and seven common wasteful patterns to avoid.