Nearly everything the agent reads or writes consumes tokens. This includes system and project instructions, your messages, earlier replies, source files, search results, command output, error logs, documentation, diffs, tool definitions, and summaries. Repeated content can be charged and processed again unless caching applies.
Large test logs, generated files, and broad searches can consume the window quickly without adding much value. Ask for focused output, open only relevant file sections, and filter noisy command results. Keep standing rules lean because they appear in many requests. When a task changes direction, start a fresh session instead of carrying unrelated history. Token awareness is not only about cost; it protects the model’s attention for the evidence that matters.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓