Skip to content

Summarization ​

As conversations grow longer, sending the entire history becomes impractical. Iris automatically condenses older messages into narrative summaries that preserve context, emotional dynamics, and unresolved threads. Summaries are scoped to individual threads — each thread maintains its own independent summary chain.

Why Summarization Matters ​

Without summarization, you'd face a tradeoff:

  • Keep all messages: Context window fills up, costs increase, responses slow down
  • Drop old messages: Lose important context, Iris forgets what you discussed

Summarization offers a middle path: older messages are compressed into rich summaries that capture what matters, while recent messages stay in full detail.

How It Works ​

Summarization now triggers based on token budget usage rather than message count. Before each LLM request, Iris checks the prompt_tokens reported by the previous turn against the model's context window:

if prompt_tokens >= context_window × compaction_threshold → summarize

At the default compaction_threshold of 0.75, summarization fires when the last request consumed 75% or more of the context window — before the window overflows, not after. The process:

  1. Check token usage: Compare the most recent turn's prompt_tokens against context_window × context.compaction_threshold
  2. Trigger if above threshold: If token usage exceeds the threshold, start summarization
  3. Select turns to summarize: Take older turns from the thread, leaving the most recent context.prune_protect_turns turns in full detail
  4. Generate summary: An LLM creates a structured summary capturing key information
  5. Mark messages as summarized: Link the summarized messages to the new summary

The protected tail (most recent turns) stays in full detail so Iris can reference recent exchanges naturally. Each thread tracks its own summarization state independently — a long-running work thread might have several summaries while a short thread about dinner plans has none.

What Gets Captured ​

Summaries aren't just text excerpts -they're structured documents that capture multiple dimensions of the conversation:

Narrative Summary ​

A 150-300 word narrative that tells the story of the conversation segment. This is what gets injected into the system prompt.

Example:

"The user discussed their ongoing Laravel project, expressing frustration with performance issues in the API layer. We explored several optimization strategies including query caching and eager loading. The conversation shifted to their upcoming vacation plans, and they mentioned needing to hand off the project to a colleague named Marcus. The user seemed stressed about the timeline but optimistic about the technical solutions we discussed."

Emotional Markers ​

Key emotional moments with intensity scores (0.0-1.0):

json
[
  {"moment": "Expressed frustration with API performance", "intensity": 0.7},
  {"moment": "Relief when caching solution clicked", "intensity": 0.6},
  {"moment": "Excitement about vacation plans", "intensity": 0.5}
]

Thread Tracking ​

Unresolved threads - topics that came up but weren't concluded:

  • "Performance testing before handoff"
  • "Meeting with Marcus about the project"

Resolved threads - topics that reached a conclusion:

  • "Caching strategy for API endpoints"
  • "Vacation dates confirmed"

Relationship Dynamics ​

How trust and rapport evolved during this segment:

  • Formality level changes
  • Building understanding
  • Areas of strong agreement or disagreement

Key Facts ​

Important information learned during this segment that might warrant memory extraction:

  • "Works with a colleague named Marcus"
  • "Has vacation planned soon"
  • "API performance is a current priority"

Summary Chaining ​

Within each thread, summaries form a chain, with each referencing its predecessor via previous_summary_id. This creates continuity within a thread's conversation history.

Thread: "Work Project"

Summary 1 → Summary 2 → Summary 3 (most recent)
    ↑           ↑           ↑
  Links to    Links to    Injected
  nothing     #1          into context

Each summary includes a narrative thread — a bridging sentence that connects to the previous summary:

"Continuing from our discussion about the API refactoring project..."

This helps Iris maintain conversational continuity within a thread even when the full history isn't available. Summary chains are completely independent between threads — a summary in one thread never references a summary from another.

When Summaries Are Used ​

Up to 3 recent summaries from the active thread are included in the system prompt for each request. They appear in the context after recalled memories, providing:

  • Historical context from earlier in the thread
  • Emotional continuity (Iris remembers how conversations felt)
  • Awareness of unresolved topics within the thread

Configuration ​

SettingDefaultDescription
context.compaction_threshold0.75Token usage fraction (of context window) that triggers summarization
context.prune_protect_turns2Recent turns kept in full detail during summarization
summarization.threshold40Secondary guard: unsummarized message count that also triggers summarization
summarization.timeout120API timeout in seconds for summary generation
summarization.modelclaude-sonnet-4-5Model used to generate summaries

Understanding the Settings ​

  • context.compaction_threshold: The primary trigger. When the previous turn's reported token usage reaches this fraction of the context window, summarization runs before the next request.
  • context.prune_protect_turns: The number of recent turns preserved verbatim during summarization. These turns are never included in a summary until a subsequent compaction cycle.
  • summarization.threshold: A secondary message-count guard. Summarization can also fire when unsummarized messages accumulate past this count, independent of token usage.

WARNING

Very aggressive summarization (low compaction threshold) may lose nuance from recent turns. The defaults balance context preservation with token efficiency.

Viewing Summaries ​

Summaries are stored in the conversation_summaries table, scoped to threads. You can explore them via:

bash
php artisan tinker
>>> $thread = User::first()->threads()->first();
>>> $thread->conversationSummaries()->latest()->first()

Each summary includes all the structured fields (emotional markers, threads, etc.) as JSON columns.

See Also ​

  • Context Management — how the token budget is calculated, how history is loaded against that budget, how tool output pruning reclaims space, and the full compaction lifecycle including progressive truncation fallback