Skip to content

Configuration ​

Iris is configured through config/iris.php. All settings have sensible defaults, but understanding what they control helps you tune the system for your needs.

How Configuration Works ​

The base configuration lives in config/iris.php. To customize without modifying core files, create config/iris-custom.php and override specific values. See Customization for the full details on merge strategies.

php
// config/iris-custom.php
return [
    'temporal' => [
        'timezone' => 'America/Chicago', // Default for all users
    ],
];

TIP

Individual users can override the default timezone in Settings > Profile.

Providers ​

Every Iris subsystem that calls an LLM has a configurable provider setting. By default, Iris uses Anthropic for chat and background tasks, OpenAI for embeddings, and ElevenLabs for text-to-speech. You can override any of these to use a different Prism provider — including local providers like Ollama.

php
// config/iris-custom.php
return [
    'agent' => [
        'provider' => 'ollama',
        'model' => 'llama3.1:8b',
    ],
];

Provider values are strings matching Prism's registered provider names (e.g., anthropic, openai, ollama, gemini, mistral, groq, xai, deepseek). See the Local Setup Guide for running Iris entirely with local models.

Agent Settings ​

These control the core chat behavior.

SettingDefaultDescription
agent.provideranthropicPrism provider for chat
agent.modelclaude-sonnet-4-5The model used for chat
agent.max_steps60Maximum tool iterations per request

What max_steps controls: When Iris uses tools, it can iterate multiple times before responding. A request like "schedule a meeting and remind me later" might use 2-3 tools. The default of 60 allows for complex multi-step tasks while preventing runaway loops.

Context Settings ​

These control what information Iris has access to when responding.

SettingDefaultDescription
context.recent_summaries3Number of conversation summaries included

How context builds: Each request includes recent messages from the active thread, recent summaries from that thread, pinned skill content, recalled memories, and calendar events. History is loaded token-budget-first — Iris pulls turns newest-first until the available token budget is exhausted.

Context Management ​

Iris uses a token-budget system to keep conversation history within model limits. After accounting for system prompts, tool definitions, and reserved headroom, it fills the remaining space with as much history as fits — starting from the most recent messages.

Token Budget Settings ​

SettingDefaultDescription
context.window200000The context window size (in tokens) of the configured model
context.token_budget_ratio0.70Fraction of the context window allocated to conversation history
context.compaction_threshold0.75Token usage fraction that triggers compaction before the next LLM request
context.prune_protect_turns2Number of recent turns whose tool outputs are shielded from compaction pruning

window: Set this to match your configured model's context window. For example, Claude Sonnet 4.6 and Opus 4.7 have 1M token context windows — set this to 1000000. The default of 200000 is safe for Haiku 4.5 and Sonnet 4.5.

token_budget_ratio: Controls how much of the context window is reserved for history. At the default of 0.70, a 200K window allocates 140K tokens for conversation messages. Raise this if you want more history; lower it if system prompts and tool definitions are large.

compaction_threshold: When actual token usage reaches this fraction of the context window, Iris compacts the context before the next request — summarizing older messages to free space. At 0.75, compaction triggers at 75% utilization. Lower this to compact more aggressively; raise it to allow the context to fill further before compacting.

prune_protect_turns: During compaction, tool call outputs from the most recent N turns are never pruned, even if they're large. This ensures the agent retains the context of its most recent actions. Increase this if agents seem to lose track of recent tool results.

You can override token_budget_ratio via environment variable:

bash
# .env
IRIS_CONTEXT_TOKEN_BUDGET_RATIO=0.80

Filesystem Tool Limits ​

These settings cap how much output filesystem tools return per call. Limiting output keeps individual responses from consuming a disproportionate share of the context window.

SettingDefaultDescription
filesystem.max_read_lines500Maximum lines returned per file read
filesystem.max_read_chars30000Maximum characters returned per file read
filesystem.max_grep_matches100Maximum matches returned per grep search

When output exceeds a limit, Iris truncates the response and saves the full result to disk. The truncation message tells the agent the saved file path so it can re-read specific sections using offset and limit parameters.

Tool Output Storage ​

Truncated tool outputs are saved to disk so agents can access the full content on demand.

SettingDefaultDescription
tool_output.storage_pathstorage/app/private/tool-output/Directory where full truncated outputs are saved
tool_output.ttl86400Seconds before saved output files are eligible for cleanup (24 hours)

Saved files are cleaned up automatically on a daily schedule. You can also run cleanup manually:

bash
php artisan iris:prune-tool-outputs

Truths Settings ​

Truths are stable, core facts about you that persist across conversations. Unlike memories that are retrieved contextually, Truths represent distilled knowledge earned through behavioral evidence.

SettingDefaultDescription
truths.max_dynamic7Maximum non-pinned Truths included per conversation
truths.percentile_threshold5Top N% of memories by access count are distillation candidates
truths.min_total_memories20Minimum memories required before distillation runs
truths.similarity_threshold0.40Minimum relevance score for a Truth to be included
truths.duplicate_threshold0.80Similarity threshold to detect duplicate Truths
truths.stale_days90Days without access before a Truth is considered stale
truths.crystallization_provideranthropicPrism provider for crystallization
truths.crystallization_modelclaude-sonnet-4-5Model for refining Truths with new evidence
truths.promotion_provideranthropicPrism provider for promotion evaluation
truths.promotion_modelclaude-sonnet-4-5Model for analyzing promotion candidates
truths.jobs_per_minute10Rate limit for distillation API calls

Pinned vs Dynamic Truths: You can pin any Truth to ensure it's always included, regardless of relevance scoring. The max_dynamic setting only limits non-pinned Truths - pinned Truths are unlimited.

Run distillation manually:

bash
php artisan iris:distill-truths --queue           # Process all users
php artisan iris:distill-truths --user=1 --queue  # Process specific user
php artisan iris:distill-truths --dry-run         # Preview without changes

Memory Settings ​

Memory settings control how Iris retrieves contextual information during conversations.

SettingDefaultDescription
memory.max_results7Maximum memories from semantic search
memory.similarity_threshold0.38Minimum similarity score to include a memory
memory.context_turns20Conversation turns used to generate search queries
memory.recall_provideranthropicPrism provider for recall query generation
memory.recall_modelclaude-sonnet-4-5Model for generating memory search queries

How memory retrieval works: When you start a conversation, Iris generates search queries based on recent messages and finds semantically similar memories. Only memories above the similarity threshold are included, keeping context focused on what's relevant.

Scoring Weights ​

Memories are ranked by a composite score:

(semantic × 0.60) + (recency × 0.25) + (frequency × 0.15)
WeightFactorWhat it means
semantic0.60How relevant to the current conversation
recency0.25Decays linearly over 90 days
frequency0.15How often the memory has been accessed

Certain memory types receive bonuses: relationships (+0.10), preferences and goals (+0.05), and recent events (+0.10).

TIP

If memories seem stale, increase the recency weight. If relevant memories aren't surfacing, increase semantic or lower similarity_threshold.

Extraction Settings ​

Extraction controls how memories are automatically created from conversations.

SettingDefaultDescription
extraction.threshold10Messages between extraction runs
extraction.max_memories6Maximum memories per extraction
extraction.history_limit50Recent conversation messages passed to the extraction LLM
extraction.timeout120API timeout in seconds
extraction.provideranthropicPrism provider for extraction
extraction.modelclaude-sonnet-4-5Model for analyzing conversations

The quality vs quantity tradeoff: A lower max_memories forces the extraction model to be selective, producing higher-quality memories. Increasing it may capture more information but risks storing trivial details.

Summarization Settings ​

Summarization compresses older messages to preserve context without consuming too many tokens.

SettingDefaultDescription
summarization.threshold40Unsummarized messages before triggering
summarization.buffer35Recent messages excluded from summarization
summarization.timeout120API timeout in seconds
summarization.provideranthropicPrism provider for summary generation
summarization.modelclaude-sonnet-4-5Model for generating summaries

How summarization triggers: Summarization is now primarily triggered by the token budget — when context utilization reaches context.compaction_threshold, Iris compacts older messages into summaries to free space. The threshold setting acts as a secondary message-count guard: if unsummarized messages exceed it, summarization also runs. During compaction, the most recent buffer messages are always preserved intact.

WARNING

Very aggressive compaction (low compaction_threshold) may lose conversational nuance. The defaults balance context preservation with token efficiency.

Consolidation Settings ​

Consolidation merges semantically similar memories into denser, more useful representations over time.

SettingDefaultDescription
consolidation.similarity_threshold0.80Minimum cosine similarity to cluster
consolidation.days_lookback3Days to look back for daily runs
consolidation.max_memories_per_run100Rate limit per consolidation run
consolidation.max_cluster_size5Max memories per cluster
consolidation.min_cluster_size2Min memories to form a cluster
consolidation.max_generation5Max consolidation generations
consolidation.timeout120API timeout in seconds
consolidation.provideranthropicPrism provider for consolidation decisions
consolidation.modelclaude-sonnet-4-5Model for merge decisions
consolidation.jobs_per_minute10Rate limit for API calls

Understanding generations: Original memories are Generation 0. When merged, the result is Generation 1. Memories can be re-consolidated up to Generation 5, allowing them to evolve as more related information is gathered. See Memory Consolidation for details.

Run consolidation manually:

bash
php artisan iris:consolidate-memories
php artisan iris:consolidate-memories --dry-run  # Preview without changes

Truth Consolidation Settings ​

Truth consolidation merges semantically similar Truths to reduce redundancy. Unlike memory consolidation which runs daily, truth consolidation runs weekly since Truths accumulate more slowly.

SettingDefaultDescription
truth_consolidation.similarity_threshold0.75Minimum cosine similarity to cluster
truth_consolidation.max_cluster_size5Max Truths per cluster
truth_consolidation.min_cluster_size2Min Truths to form a cluster
truth_consolidation.timeout120API timeout in seconds
truth_consolidation.provideranthropicPrism provider for truth consolidation
truth_consolidation.modelclaude-sonnet-4-5Model for merge decisions
truth_consolidation.jobs_per_minute10Rate limit for API calls

Protection model: Only Promoted Truths are eligible for consolidation. User-created, agent-created, and pinned Truths are protected. Truths at the maximum generation are also excluded.

The threshold (0.75) is lower than memory consolidation (0.80) because Truths are already distilled facts - semantic overlap is more likely to be true redundancy.

Run consolidation manually:

bash
php artisan iris:consolidate-truths
php artisan iris:consolidate-truths --dry-run  # Preview without changes

See Truth Consolidation for details on how it works.

Calendar Settings ​

SettingDefaultDescription
calendar.cache_ttl15Cache duration in minutes
calendar.event_horizon7Days ahead to include in context

Why caching matters: Calendar data is fetched from Google's API and cached to avoid repeated requests. The event_horizon controls how far ahead Iris looks -a larger value gives more scheduling context but increases the data in each request.

Embeddings Settings ​

SettingDefaultDescription
embeddings.provideropenaiPrism provider for embeddings
embeddings.modeltext-embedding-3-smallEmbedding model
embeddings.provider_options[]Provider-specific options passed to Prism

Embeddings power semantic search. The default text-embedding-3-small model offers a good balance of quality and cost. Changing the model or provider affects how memories are stored and retrieved.

The provider_options array lets you pass provider-specific configuration. For example, when using Ollama for embeddings you'll want to set dimensions to match the default vector size:

php
'embeddings' => [
    'provider' => 'ollama',
    'model' => 'your-embedding-model',
    'provider_options' => [
        'dimensions' => 1536,
    ],
],

Text-to-Speech Settings ​

NOTE

Text-to-speech is disabled by default. Enable it by setting IRIS_TTS_ENABLED=true in your .env.

SettingDefaultDescription
tts.enabledfalseEnable TTS audio playback
tts.providerelevenlabsPrism provider for text-to-speech
tts.modeleleven_multilingual_v2TTS model
tts.voice(configured)Voice ID for speech generation

When enabled, assistant messages display a play button that generates audio using ElevenLabs. See Text-to-Speech for setup instructions.

Telegram Notification Settings ​

NOTE

Telegram notifications are disabled by default. Enable them by setting TELEGRAM_ENABLED=true in your .env and configuring a bot token.

SettingEnvironment VariableDescription
connectors.telegram.enabledTELEGRAM_ENABLEDEnable Telegram integration
connectors.telegram.tokenTELEGRAM_BOT_TOKENBot token from BotFather
connectors.telegram.bot_usernameTELEGRAM_BOT_USERNAMEBot username for generating connect links

These settings live in config/connectors.php. When enabled, proactive messages are pushed to users who've linked their Telegram account and toggled notifications on. See Telegram Notifications for the full setup guide.

Thread Briefs Settings ​

Thread Briefs give Iris cross-thread awareness by generating lightweight digests of each thread's current state.

SettingDefaultDescription
briefs.enabledtrueMaster toggle for all brief-related behavior
briefs.frequency2Assistant responses between brief generations
briefs.threshold2Minimum assistant responses before first brief
briefs.max_threads5Maximum threads shown in cross-thread context
briefs.provideranthropicPrism provider for brief generation
briefs.modelclaude-haiku-4-5Model for generating briefs
briefs.timeout15Timeout in seconds for generation

How it works: After every N assistant responses (default: 2), a background job generates a 3-6 sentence digest covering the thread's topic, key decisions, current status, and emotional context. These briefs are injected into other threads' system prompts and used by the heartbeat for context. The master enabled toggle controls everything — generation, cross-thread injection, and heartbeat usage.

Disable via environment variable:

bash
# .env
IRIS_BRIEFS_ENABLED=false

Customizing Briefs ​

php
// config/iris-custom.php
return [
    'briefs' => [
        'frequency' => 3,      // Generate less often
        'max_threads' => 3,    // Fewer threads in cross-thread context
    ],
];

Proactive Messages Settings ​

SettingDefaultDescription
heartbeat.provideranthropicPrism provider for heartbeat decisions
heartbeat.modelclaude-sonnet-4-5Model for heartbeat decisions
heartbeat.max_steps30Maximum tool iterations when crafting messages
heartbeat.history_limitnullConversation messages included (null = use default)
heartbeat.context_max_threads3Most recently active threads to include in heartbeat context
heartbeat.prompts(see below)Prompt stack for heartbeat system messages

How it works: Iris runs a heartbeat every 30 minutes, gathering context (thread briefs and latest summaries from active threads, thread metadata, calendar, weather, memories) and deciding whether to proactively reach out. The heartbeat uses a dedicated prompt stack that's lighter than the full conversation prompt stack. Users configure guidance (soft preferences) and boundaries (hard constraints like quiet hours) through the UI.

See Proactive Messages for detailed usage and configuration.

Temporal Settings ​

SettingDefaultDescription
temporal.timezoneAmerica/New_YorkDefault timezone for date/time context

This timezone is injected into the system prompt so Iris knows the current date and time. It serves as the system-wide default for all users.

Per-User Timezone ​

Each user can override the default timezone in Settings > Profile. When a user sets their timezone, it's used everywhere: system prompts, heartbeat scheduling, quiet hours, follow-up validation, and the frontend.

The fallback chain is:

  1. User's timezone (if set in their profile)
  2. Config default (temporal.timezone)

If a user leaves timezone blank, the config default applies automatically. This makes the config value a sensible default for new users while letting individuals customize.

Shell Settings ​

WARNING

Shell command execution is disabled by default. Enable only on trusted deployments.

SettingDefaultDescription
shell.enabledfalseEnable shell command execution
shell.default_timeout30Default timeout in seconds
shell.max_timeout300Maximum allowed timeout (5 minutes)
shell.max_output_length50000Maximum output size in bytes
shell.default_working_directorynullDefault directory for commands
shell.blocked_executables['sudo', ...]Blocked command names
shell.blocked_patterns[...]Regex patterns to block
shell.inherit_env_vars['PATH', ...]Environment variables to inherit

Enable via environment variable:

bash
# .env
IRIS_SHELL_ENABLED=true

Security layers: Privilege escalation commands are blocked (sudo, su, doas, pkexec). Dangerous patterns like rm -rf / and direct disk writes are detected and blocked. Only safe environment variables are inherited.

See Shell Commands for detailed usage and security information.

Broadcasting Settings ​

Iris uses Laravel Reverb for real-time streaming. Configure via environment variables:

SettingDefaultDescription
BROADCAST_CONNECTIONreverbBroadcasting driver
REVERB_APP_ID-Application identifier
REVERB_APP_KEY-Client authentication key
REVERB_APP_SECRET-Server-side secret
REVERB_HOSTlocalhostReverb server hostname
REVERB_PORT8080Reverb server port
REVERB_SCHEMEhttpProtocol (http or https for production)

Event Storage ​

Stream events are temporarily stored in Redis for connection recovery. Events remain available for replay, allowing clients to catch up after brief disconnections.

Sub-Agent Settings ​

WARNING

Task delegation is disabled by default. Requires shell commands to also be enabled.

SettingDefaultDescription
subagent.enabledfalseEnable task delegation
subagent.provideranthropicPrism provider for the sub-agent
subagent.modelclaude-sonnet-4-5Model for the sub-agent
subagent.max_steps30Maximum tool iterations per task
subagent.default_working_directorynullDefault directory for tasks
subagent.request_timeout120HTTP timeout per API request
subagent.default_timeout300Default task timeout (5 minutes)
subagent.max_timeout600Maximum task timeout (10 minutes)

Enable via environment variables:

bash
# .env
IRIS_SUBAGENT_ENABLED=true
IRIS_SHELL_ENABLED=true
IRIS_SUBAGENT_WORKING_DIR=/home/user/projects  # Optional default directory

How it works: Task delegation spawns an autonomous sub-agent using Prism's agent loop. The sub-agent has access to shell commands and web tools, executing multi-step tasks independently before returning results. Tasks are processed by the iris:agent daemon — see Task Delegation for setup.

Constraints: One active task per user to prevent resource exhaustion. Inherits all shell security settings.

See Task Delegation for detailed usage information.

Token Usage Tracking ​

Iris tracks API token consumption across all operations. You can monitor usage by source type:

SourceWhat it tracks
chatMain conversation tokens
summarizationSummary generation tokens
extractionMemory extraction tokens
consolidationMemory consolidation tokens
truth_consolidationTruth consolidation tokens
promotionTruth promotion analysis tokens
crystallizationTruth crystallization tokens
embeddingEmbedding API calls

Token data is stored in the token_usages table and can be queried for cost analysis or monitoring dashboards.

Common Configuration Scenarios ​

Lower API Costs ​

php
// config/iris-custom.php
return [
    'extraction' => [
        'threshold' => 20,  // Extract less frequently
    ],
    'summarization' => [
        'threshold' => 60,  // Summarize less often
    ],
    'truths' => [
        'max_dynamic' => 3,  // Fewer dynamic Truths per conversation
    ],
    'memory' => [
        'max_results' => 5,  // Fewer search results
    ],
];

More Responsive Memory ​

php
// config/iris-custom.php
return [
    'extraction' => [
        'threshold' => 5,    // Extract more frequently
    ],
    'memory' => [
        'max_results' => 10,              // More search results
        'similarity_threshold' => 0.10,   // Lower threshold
    ],
];

Conservative Memory Consolidation ​

php
// config/iris-custom.php
return [
    'consolidation' => [
        'similarity_threshold' => 0.90,  // Only very similar memories
        'max_generation' => 3,           // Limit consolidation depth
    ],
];

Conservative Truth Consolidation ​

php
// config/iris-custom.php
return [
    'truth_consolidation' => [
        'similarity_threshold' => 0.85,  // Only very similar truths
    ],
];

Aggressive Truth Promotion ​

php
// config/iris-custom.php
return [
    'truths' => [
        'percentile_threshold' => 20,    // Wider candidate pool
        'min_total_memories' => 10,      // Start promoting sooner
    ],
];

Enable Shell Commands ​

php
// config/iris-custom.php
return [
    'shell' => [
        'default_timeout' => 60,
        'default_working_directory' => '/home/user/projects',
    ],
];

Then set IRIS_SHELL_ENABLED=true in your .env.

Enable Task Delegation ​

php
// config/iris-custom.php
return [
    'subagent' => [
        'max_steps' => 50,
        'default_timeout' => 600,  // 10 minutes for complex tasks
        'default_working_directory' => '/home/user/projects',
    ],
];

Then set both IRIS_SUBAGENT_ENABLED=true and IRIS_SHELL_ENABLED=true in your .env.

Optimize for 200K Models ​

For models with 200K context windows (Claude Haiku 4.5, Claude Sonnet 4.5), keep context lean to avoid hitting limits during long conversations.

php
// config/iris-custom.php
return [
    'context' => [
        'window' => 200_000,
        'token_budget_ratio' => 0.65,   // Leave more headroom for system prompts
        'compaction_threshold' => 0.70, // Compact earlier to avoid hitting limits
        'prune_protect_turns' => 2,
    ],
    'filesystem' => [
        'max_read_lines' => 300,        // Smaller reads to conserve context
        'max_read_chars' => 15000,
        'max_grep_matches' => 50,
    ],
];

Maximize Context on 1M Models ​

For models with 1M context windows (Claude Sonnet 4.6, Claude Opus 4.7), you can afford much larger history and tool outputs.

php
// config/iris-custom.php
return [
    'context' => [
        'window' => 1_000_000,
        'token_budget_ratio' => 0.80,   // Use more of the window for history
        'compaction_threshold' => 0.85, // Allow context to fill further before compacting
        'prune_protect_turns' => 5,     // Protect more recent tool outputs
    ],
    'filesystem' => [
        'max_read_lines' => 2000,       // Read much larger files in one shot
        'max_read_chars' => 100000,
        'max_grep_matches' => 500,
    ],
];

Thread Settings ​

Threads organize conversations into isolated contexts with independent history and summaries.

SettingDefaultDescription
threads.naming_threshold4Messages before auto-naming triggers
threads.naming_provideranthropicPrism provider for naming
threads.naming_modelclaude-haiku-4-5Model for generating thread names
threads.naming_timeout15Timeout in seconds for the naming job

How auto-naming works: After a new thread accumulates enough messages (default: 4), a background job generates a short, descriptive name. A fast, inexpensive model is used to keep costs minimal. If a user manually renames a thread, auto-naming won't overwrite their choice.

Customizing Thread Naming ​

php
// config/iris-custom.php
return [
    'threads' => [
        'naming_threshold' => 6,  // Wait for more conversation context
    ],
];

Skills Settings ​

SettingDefaultDescription
skills.enabledtrueEnable agent skills
skills.directory.agents/skillsDirectory to load skills from

Skills extend Iris with specialized knowledge and workflows. You can pin skills to threads to inject their full content into the system prompt for every message in that thread. Pinned skills are configured per-thread through the thread settings modal. See Agent Skills for details.