Skip to content

Chat Interface ​

The chat interface orchestrates the full conversation flow — assembling context, streaming responses, executing tools, and persisting messages. Understanding this flow helps you see how all of Iris's components work together.

Every conversation happens within a thread. Threads provide isolation — conversation history, summaries, streaming, and pinned skills are all scoped to the active thread.

Request Flow ​

When you send a message, Iris processes it through several stages before streaming a response:

1. Context Recall ​

The ContextRetriever fetches relevant context using two complementary systems:

  • Truths: Stable, core facts about you (pinned Truths always included, others ranked by relevance)
  • Memories: Semantic search finds memories related to the current conversation

This happens first because context influences how Iris responds.

2. Context Assembly ​

Multiple context sources are gathered in parallel:

  • Recent conversation history from this thread (up to 50 messages)
  • Conversation summaries from this thread (up to 3 recent summaries)
  • Thread briefs from other active threads (for cross-thread awareness)
  • Pinned skill content for this thread (if any)
  • Calendar events (next 7 days, if connected)
  • The current date and time

3. System Prompt Building ​

The SystemPromptBuilder assembles a personalized prompt by rendering each registered prompt class in order. The result includes Iris's identity, recalled memories, summaries, calendar context, and temporal information.

4. Process and Stream ​

The request is queued for asynchronous processing via Prism PHP. Response events stream to your browser via WebSockets in real-time, including text chunks and tool calls. If Iris decides to use a tool, execution happens server-side before continuing the response.

Stream events are scoped to the active thread — only clients viewing that thread receive the broadcast. This prevents messages from appearing in the wrong thread when you have multiple tabs open.

This architecture provides reliable delivery with automatic recovery from connection drops.

5. Persist and Process ​

After the response completes:

  • The conversation is saved to the database within the active thread
  • Token usage is recorded for monitoring
  • Background jobs are dispatched (memory extraction, thread-scoped summarization)

Agentic Behavior ​

Iris operates as an agent, meaning it can use tools and iterate multiple times before providing a final response. The agent.max_steps configuration (default: 60) limits how many iterations can occur.

This enables complex, multi-step tasks:

  • Search and ask: Search memories, find nothing relevant, then ask a clarifying question
  • Create and remember: Create a calendar event, then store a memory about why it was scheduled
  • Generate and describe: Generate an image, then describe what was created

Each tool invocation is a step. A simple memory storage is 1 step. A complex request might use 5-10 steps as Iris searches, creates, and confirms.

NOTE

Tool calls stream to the frontend in real-time, so users see what Iris is doing as it works.

Features ​

Text-to-Speech ​

Assistant messages include a play button that reads the message aloud using ElevenLabs. Click once to play, click again to stop. This feature requires setup - see Text-to-Speech for configuration.

Image Support ​

Users can upload images for multi-modal conversations. Images are attached to messages and sent to Claude as part of the request. Common use cases include:

  • Asking questions about screenshots
  • Getting feedback on designs
  • Extracting information from photos

Retry ​

If a provider error occurs, the failed message appears inline with a retry button. Clicking retry re-runs the full request — context assembly, LLM call, and tool execution — with the same input. See Error Handling for details on what information is shown and when retrying helps.

Streaming ​

All responses stream in real-time via WebSockets using Laravel Reverb. Streams are scoped to the active thread, so each thread's broadcasts are isolated. The stream includes:

Event TypeContent
Text chunksPartial response text as it's generated
Tool callsWhen Iris invokes a tool
Tool resultsWhat the tool returned
Provider toolsActivity from Anthropic's built-in tools
ArtifactsGenerated content like images

Connection Recovery ​

If your connection drops briefly, you won't miss any events:

  • Events are temporarily stored in Redis for replay
  • Clients automatically catch up when reconnecting
  • Sequence numbers ensure no duplicates

Stopping Streams ​

Users can stop a stream mid-generation. The stop is graceful—partial responses are preserved and the conversation state remains consistent.

Request Architecture ​

When you send a message, Iris processes it asynchronously:

  1. Accept & Queue: Your message is received within the active thread and a background job is dispatched
  2. Process: The job builds thread-scoped context, executes the agent, and generates a response
  3. Broadcast: Response events stream to clients viewing this thread via WebSockets
  4. Persist: The conversation is saved to the thread and background jobs (extraction, summarization) are queued

This architecture enables:

  • Reliable delivery: If your connection drops, events are stored and replayed when you reconnect
  • Fail-fast errors: Provider errors surface immediately with no automatic retries — the error appears inline with a retry button
  • Non-blocking responses: Your browser remains responsive while processing happens server-side

Error Handling ​

When a provider request fails, Iris does not retry automatically — the job runs once (tries = 1), and if it fails, the error surfaces immediately. Iris marks the conversation as failed and broadcasts the error to your browser.

What you see ​

The failed assistant message appears inline in the chat with two pieces of information:

  • Human-readable message: A plain-English description of what went wrong
  • Retry button: Always available — all errors are currently marked as recoverable

Error messages by type ​

ErrorMessage shown
Prompt too long (context window exceeded)"The conversation is too long for the model's context window. Try starting a new conversation or asking Iris to work in smaller steps."
Rate limited"The AI provider is currently rate-limited. Please wait a moment and try again."
Provider overloaded"The AI provider is temporarily overloaded. Please try again in a moment."
Other provider errorsThe provider's own message, or a generic fallback

Retrying manually ​

Click the retry button below a failed message to re-run the full request with the same input. For transient errors (rate limits, provider overload), retrying immediately or after a brief wait usually resolves the issue.

Prompt too long ​

If Iris reports that the conversation is too long for the model's context window, retrying with the same input won't help — the problem is the accumulated context size, not a transient provider state. Your options are:

  • Start a new conversation in a fresh thread with no history
  • Ask Iris to work in smaller steps — break the task into pieces that each fit within the window

NOTE

Context management (compaction, pruning, and progressive truncation) keeps most conversations well within bounds during normal use. See Context Management for how Iris handles a filling context window before this limit is reached.

Failed conversations are preserved ​

Failed conversations are never deleted. The message stays in the chat with its error state and the retry button visible — even after navigating away and returning. Iris retains the full failed state so you always see what was attempted and what went wrong.

Frontend Integration ​

The frontend uses Laravel Echo for WebSocket connections. A custom hook manages connection state, event sequencing, and automatic replay on reconnection.