Streaming Responses
MindRoom streams agent responses to Matrix by progressively editing a single message, so users see text appear as the model generates it. This page covers when streaming is used, how to tune or disable it, how streamed messages show tool calls, completion, cancellation, and errors, and how oversized responses are delivered.
How It Works
- MindRoom sends an initial message showing a
Thinking...placeholder until text arrives. - As text arrives, MindRoom edits the same message with the accumulated content.
- When the response finishes, a final edit holds the complete text.
When a response is not streamed, the placeholder is replaced with the complete response in a single edit. Agents also show a typing indicator while they work on a response.
Configuration
Streaming is enabled by default.
These settings are global-only under defaults and cannot be overridden per agent.
defaults:
enable_streaming: true # Set false to disable progressive edits
streaming:
update_interval: 5.0
min_update_interval: 0.5
interval_ramp_seconds: 15.0
max_idle: 2.0
| Field | Type | Default | Description |
|---|---|---|---|
enable_streaming |
bool | true |
Stream responses via progressive message edits. When false, the placeholder is replaced with the complete response in a single edit. |
streaming.update_interval |
float, > 0 | 5.0 |
Steady-state seconds between edits. |
streaming.min_update_interval |
float, > 0 | 0.5 |
Seconds between edits at the start of a response. |
streaming.interval_ramp_seconds |
float, >= 0 | 15.0 |
Seconds over which the interval ramps from min_update_interval to update_interval. Set 0 to use update_interval from the start. |
streaming.max_idle |
float, > 0 | 2.0 |
When no new text has arrived for this many seconds, the buffered text is shown with the next streaming update. |
Edits also come sooner than the interval once enough new text has accumulated, and a tool-call marker appears as soon as the tool starts.
Raise update_interval to reduce load on the homeserver, or lower it for smoother updates.
Presence-Based Streaming
Even when streaming is enabled, MindRoom streams a response only when the user who sent the message is online or unavailable; for an offline requester, the placeholder is replaced with the complete response.
Everyone in the room sees the same edits.
If the presence check fails, for example because the homeserver has presence disabled, MindRoom does not stream.
A reply that continues after a restart always streams, even when streaming is off or the requester is offline, so the part it already showed stays in the message.
When no requester can be identified, MindRoom streams.
Presence
Running agents appear online (or unavailable when idle) and offline when stopped or disconnected.
Their presence status message shows the model, role, and tool count, for example 🤖 Model: anthropic/claude-sonnet-5-5 | 💼 Code assistant | 🔧 5 tools available.
Stream Status
Each streamed message carries its progress in the Matrix event content field io.mindroom.stream_status.
| Value | Meaning |
|---|---|
pending |
Initial message of the response |
streaming |
Response is still being generated |
approval_pending |
Response is paused until someone approves or denies a tool call (see Tool Approval) |
completed |
Response finished successfully |
cancelled |
User stopped the response |
error |
Response failed or was otherwise interrupted |
To confirm a response finished, inspect the latest io.mindroom.stream_status in a client or event view that shows event content.
MindRoom never appends an ellipsis to generated text, and clients may draw their own progress indicators, so a trailing ellipsis is not a progress signal.
Tool Calls During Streaming
When an agent calls tools, MindRoom shows inline markers in the message text:
The number in brackets counts tool calls within the message, and io.mindroom.tool_trace.events[N-1] in the event content holds the tool name, argument preview, and result preview for call [N].
Tools that start in quick succession appear together in one edit.
Markers and tool-trace metadata are shown by default; to hide them, set show_tool_calls: false on one agent or under defaults for every agent that leaves it unset (see Agent Configuration).
With tool calls hidden, the agent still shows typing activity, and a tool that needs an isolated worker may show generic progress such as Preparing isolated worker... 17s elapsed. without naming the tool.
Cancellation and Errors
Users can cancel an in-progress response by reacting with 🛑 on the message being generated (see Stop Button).
Tool calls running in a worker, such as run_shell_command or run_python_code, stop with it, apart from the exceptions listed there.
The partial text stays in the message, followed by a note explaining how it ended:
| Note | Cause |
|---|---|
**[Response cancelled by user]** |
The user stopped the response. |
**[Response interrupted by service restart]** |
MindRoom restarted while the response was being generated; the reply continues below the note once MindRoom is back, or ends there if nothing is left to continue. |
**[Response interrupted]** |
The response was interrupted for another reason. |
**[Response interrupted by an error: <error description>]** |
Generation failed; the description says why. |
A crash, an orderly shutdown, or an agent restarted by a configuration or MCP change, including during an approved tool call, keeps the partial reply and its tool calls, and the agent continues it in the same message once MindRoom is back.
The agent is told which visible tool calls already finished so it does not repeat side effects; calls hidden by show_tool_calls: false cannot be passed on, so it is only warned that they may have run.
Large Messages
A Matrix event holds at most about 64 KB, so defaults.large_message_strategy controls how a larger response is delivered.
| Value | Delivery |
|---|---|
sidecar (default) |
A preview with as much text as fits, plus the full message content attached as message-content.json (message-content.json.enc, encrypted, in encrypted rooms). |
split |
Several complete rich-text messages: the first replaces the streaming placeholder and the rest follow it in the same thread or room. |
The sidecar strategy applies to messages over about 55 KB and edits over about 27 KB, and the preview omits inline tool-trace metadata.
With split, each segment renders as standalone Markdown, text is cut at paragraph or line boundaries and never inside a fenced code block, and the segments together reproduce the full response exactly.
Even with split, non-text payloads, metadata that alone exceeds the size budget, and a single code block larger than one event use the sidecar.