Skip to content

Context Management ​

Every request Iris sends to an LLM has a fixed token limit — the model's context window. Context management is the set of mechanisms Iris uses to keep that request within bounds while preserving as much relevant history as possible.

This page covers how the token budget is calculated, how conversation history is loaded against that budget, how tool output pruning reclaims space invisibly, and how compaction summarizes old turns when the window fills up.

Token Budget Calculation ​

Before loading any history, Iris calculates how many tokens are available. The formula is:

available tokens = (context window × token_budget_ratio) - system prompt tokens - tool definition tokens

The token estimate uses a chars/4 heuristic: ceil(strlen($text) / 4). This isn't perfect, but it's fast and errs conservatively.

Config keys ​

KeyDefaultEffect
context.window200000The context window size (in tokens) of your configured model
context.token_budget_ratio0.70Fraction of the context window allocated to conversation history

Set context.window to match the model you're using. For example, Claude Sonnet 4.6 and Opus 4.7 have 1M token context windows — set context.window to 1000000 to take advantage of the full capacity. The default of 200K is safe for models like Haiku 4.5 and Sonnet 4.5.

On a 200K window, after accounting for a typical system prompt and tool definitions (~10K tokens combined), you'll have ~130K tokens for conversation history. On a 1M window, that grows to ~690K tokens — roughly 5× more conversation depth.

TIP

To allow more history at the cost of less headroom for responses, raise context.token_budget_ratio to 0.80 or 0.85. This can help on smaller context windows during long tool-heavy sessions.

Budget-Based History Loading ​

Iris loads conversation history starting from the most recent turn and works backwards, accumulating turns until the token budget is exhausted.

Unlike a message-count approach, this treats turns by size, not by number. A conversation with 50 long tool outputs and 50 short exchanges would consume wildly different amounts of context — the token-budget approach accounts for that. The model always sees the most recent turns first, never truncating from the middle.

How history loading works:

  1. Query all conversations from the active thread, newest-first
  2. Calculate available token budget using TokenBudgetCalculator
  3. Iterate through conversations, accumulating token estimates, until the budget is exhausted
  4. Reverse the selected messages to chronological order before passing to the LLM

Config keys ​

KeyDefaultEffect
context.token_budget_ratio0.70Controls how many tokens are available for history

Tool Output Pruning ​

Even with budget-based loading, tool outputs from older turns can waste tokens. A read_file call that returned 5,000 characters of file content is useful context when it happens — but ten turns later, the agent probably doesn't need all of that content verbatim.

Pruning replaces the content of old tool results with a placeholder ([Tool output cleared]) in the LLM request. It happens at request-build time, not as a background job, and it's completely invisible to you: full content is always retained in the database.

What gets pruned ​

Every ToolResultMessage outside the most recent N turns is a pruning candidate. The N most recent turns are always protected.

A turn is one user message plus all subsequent assistant messages and tool result messages until the next user message. With context.prune_protect_turns = 2, the two most recent turns are shielded — their tool outputs are sent verbatim. All older turns have their tool outputs replaced with the placeholder.

Config keys ​

KeyDefaultEffect
context.prune_protect_turns2Number of recent turns whose tool outputs are never pruned

The invariant: database content is always preserved ​

Pruning operates exclusively on in-memory Prism message objects. It never modifies the conversations table. You can always inspect the full tool output of any past turn directly in the database — pruning only affects what's sent to the LLM for a given request.

IMPORTANT

Full conversation content — including all tool outputs — is always retained in the database regardless of pruning or compaction state.

Compaction ​

When the context window fills up despite pruning, Iris compacts the conversation. Compaction summarizes older turns into a structured narrative, freeing the window for new activity.

When compaction triggers ​

Before each LLM request, Iris checks the prompt_tokens from the most recent completed turn against the model's context window:

if prompt_tokens >= context_window × compaction_threshold → compact

At the default compaction_threshold of 0.75, compaction triggers when the last request used 75% or more of the model's context window. This means compaction fires before the window overflows, not after.

What the summary includes ​

Compaction generates a structured ConversationSummary with these fields:

FieldDescription
summary150–300 word narrative of the conversation segment
accomplishmentsCompleted tasks and outcomes from this segment
key_decisionsChoices made and conclusions reached
relevant_filesFiles touched, modified, or referenced
active_goalsOpen goals or tasks still in progress
emotional_threadHow the emotional tone evolved through this segment
relationship_dynamicsShifts in formality, rapport, and trust
evolving_themesTopics that developed or transformed across turns

Recent turns are kept verbatim — compaction only affects the turns being summarized, not the tail that will still appear in full.

For a deeper look at what summaries capture and how the summary chain works, see Summarization.

Config keys ​

KeyDefaultEffect
context.compaction_threshold0.75Token usage fraction that triggers compaction
context.prune_protect_turns2Recent turns kept verbatim during compaction

Progressive truncation fallback ​

If the compacted summary plus protected recent turns still exceed the context window, Iris applies progressive fallback:

  1. Reduce the protected turn window step-by-step (configured value → 1 → 0)
  2. If still over budget with all tool outputs pruned, throw ConversationTooLargeException

This is a last resort — normal compaction almost never reaches step 2.

Walkthrough: A Long Tool-Heavy Conversation ​

Here's how the full lifecycle plays out in practice.

Setup: You're running on a 200K context model. System prompt + tool definitions occupy ~10K tokens. The budget ratio is 0.70, so ~130K tokens are available for history.


Turns 1–5: Early conversation, tools used freely

The budget check shows 130K tokens available. History loading grabs all 5 turns — they fit easily. Pruning protects the 2 most recent turns; turns 1–3 have their tool outputs replaced with [Tool output cleared]. But this is invisible: turns 1–3 still appear in the database in full.


Turns 6–20: Tool-heavy accumulation

You ask Iris to explore a large codebase — multiple read_file calls, grep searches, shell commands. Each turn produces thousands of characters of tool output. History loading can still fit all 20 turns because pruning shrinks the effective token cost of older turns. The protected window (last 2 turns) always has full tool output; everything older is a placeholder.


Turn 21: Compaction triggers

The prompt_tokens reported back from the last request is 158K — above the 75% threshold of 150K. Before sending turn 21 to the LLM, Iris runs the Summarizer on the oldest turns not yet summarized. The result is a ConversationSummary covering turns 1–18, capturing accomplishments (files explored, patterns found), relevant files (the ones you looked at), active goals (the refactoring task still in progress), and emotional thread (curiosity, some frustration at a confusing module).

Turns 19–21 are kept in full detail. The LLM now sees:

  • The new summary injected via the system prompt
  • Turns 19–21 in full (with pruning applied to 19's tool outputs)
  • Your new message (turn 21)

Turn 22 onward: Conversation continues

The context is spacious again. History loading now fits turns 19–21 plus any new turns. As the conversation grows, the cycle repeats: pruning reclaims token space turn by turn, and compaction fires again when utilization creeps back up toward 75%.


What you see: Nothing. The conversation flows naturally. Iris remembers the files you explored and the goals you set even after they've been summarized away from the raw message list.

Cross-References ​

  • Summarization — what conversation summaries capture and how the summary chain works
  • Configuration: Context Management — full reference for all config keys: context.token_budget_ratio, context.compaction_threshold, context.prune_protect_turns