Voice Calls
MindRoom agents can join Element Call voice calls in their rooms and talk with you in real time.
The realtime backend provides OpenAI realtime speech-to-speech.
The live backend uses OpenAI Live for speech and delegates substantive requests to the normal MindRoom agent.
The cascaded backend combines independently configured speech-to-text and text-to-speech services with the agent's normal MindRoom response path.
How it works
Matrix group calls (MatrixRTC, used by Element Call, Element X, and recent Cinny releases) do not send media over Matrix itself.
Matrix only carries the signaling: participants publish org.matrix.msc3401.call.member state events, and media flows through a LiveKit SFU that is deployed next to the homeserver.
When a call starts in a room, the configured agent:
- Sees the call membership state event and re-reads the room state.
- Exchanges a Matrix OpenID token for a LiveKit JWT at the MatrixRTC authorization service (
lk-jwt-service). - Connects to the LiveKit SFU and publishes its own call membership state event, so it appears in the call roster.
- In encrypted rooms, distributes its media frame key over encrypted to-device messages and installs the other participants' keys, following the same per-sender key rotation policy as Element Call.
- Runs the selected voice backend until the managed session ends.
The voice agent is the same agent you chat with.
Realtime carries the agent's rendered prompt and effective tools into OpenAI Realtime.
Live receives the normal agent's full rendered system prompt, including its configured context files and caller-specific workspace, followed by voice delivery and delegation guidance. It can answer from that context directly and delegates requests needing tools, research, additional memory, or actions to the normal MindRoom agent. Tool schemas and execution remain in the delegated agent. The prompt is prepared before the first greeting, and the prepared agent is reused for delegated turns.
The voice model relays the agent's results conversationally.
Cascaded sends each finalized transcript through the normal MindRoom agent response path, preserving model resolution, the rendered system prompt, instructions, knowledge, skills, hooks, tools, requester identity, history storage, and tool execution behavior.
A cascaded profile may explicitly override the resolved LLM while preserving all other agent behavior.
MindRoom keeps a managed call active only while its devices belong to one Matrix user, and uses that user's ID as the real requester for the room-scoped tool runtime.
Multiple devices belonging to that same user are supported.
Tools requiring confirmation, user input, external execution, or tool_approval are hidden in all backends because a voice call has no approval UI.
The agent leaves the call and clears its membership state event when the caller leaves, a second distinct Matrix user joins, the voice session terminates, or the bot shuts down.
Configuration
calls:
enabled: true
profiles:
openai-realtime:
backend: realtime
model: gpt-realtime-2.1
credentials_service: openai-voice
voice: marin
agents:
assistant: openai-realtime
# livekit_service_url: https://rtc.example.org # same-server .well-known override
Voice calls require the matrix_calls extra (pip install "mindroom[matrix_calls]" or uv sync --extra matrix_calls).
calls.profiles defines reusable, complete voice pipelines.
calls.agents maps each enabled agent to exactly one profile name.
The model in a realtime or Live profile is the provider-specific OpenAI speech model ID.
Live profiles require an explicit credentials_service and voice, and do not use separate STT or TTS services.
The optional agent_model in a Live profile references a named entry from the top-level models mapping for delegated agent turns.
Cascaded profiles require complete STT and TTS settings.
OpenAI speech legs require either credentials_service or api_key; OpenAI-compatible speech legs require host and may omit credentials.
The optional model in a cascaded profile references a named entry from the top-level models mapping.
An explicit cascaded call model takes precedence over room and agent models for the calls-enabled agent.
Omitting it keeps the existing room and agent model resolution.
There are no inherited backend defaults or credential fallbacks.
enabled and livekit_service_url remain global and cannot be overridden per agent.
Set a dedicated credential service when voice and chat use different OpenAI credentials.
A missing selected service does not fall back to another key.
For example, these two agents use independent voice pipelines:
calls:
enabled: true
profiles:
openai-realtime:
backend: realtime
model: gpt-realtime-2.1
credentials_service: openai-voice
voice: marin
local-cascaded:
backend: cascaded
stt:
provider: openai_compatible
model: whisper-large-v3
host: http://127.0.0.1:9000
tts:
provider: openai_compatible
model: kokoro
host: http://127.0.0.1:9001
extra_kwargs:
voice: af_heart
agents:
concierge: openai-realtime
local_assistant: local-cascaded
MindRoom enforces at most one calls-enabled agent per room.
Calls only join rooms configured for that agent and only while the sole caller passes the normal room and per-agent reply permissions.
Calls-enabled agents also join calls in ad-hoc rooms they accepted through their normal authorized-invite policy. This lets Matrix clients create a private, temporary voice room and invite one agent without adding that room to config.yaml first.
An explicit room assignment to one calls-enabled agent takes precedence over persisted invite records; invite-only ambiguity still fails closed.
Calls-enabled agents advertise 📞 Voice calls in their Matrix presence status only when their MatrixRTC runtime is available, so clients can show the call action only where it will be answered.
Requester-private agents use the sole authorized caller's verified Matrix user ID to resolve their workspace, memory, credentials, history, knowledge, and tool execution scope.
The same caller scope applies in configured rooms and authorized ad-hoc invite rooms.
OpenAI Live with agent delegation
This example uses GPT-Live for speech and the configured chat model for the normal agent's delegated responses.
The voice model uses its own named OpenAI credential service, independently of the agent's model provider and credentials.
models:
chat:
provider: anthropic
id: claude-opus-5
agents:
assistant:
display_name: Assistant
model: chat
calls:
enabled: true
profiles:
openai-live:
backend: live
model: gpt-live-1
credentials_service: openai-live
voice: marin
agent_model: chat
agents:
assistant: openai-live
agent_model is optional; omitting it keeps normal room and agent model resolution.
When set, it overrides that resolution for the calls-enabled agent and must name an entry in models.
An unknown alias fails configuration loading.
The model field selects the OpenAI Live speech model and is independent of agent_model.
The named credentials_service must resolve an API key; a missing service does not fall back to another OpenAI key.
Live needs no separate STT or TTS credentials.
Media reconnects reuse the delegate's history for the current call; once the caller leaves, the next call starts a new delegate session.
Configuration reload rebuilds the call runtime when delegate model definitions or room model routing change.
Cascaded LLM model override
This example uses a large model for text conversations and a faster model for cascaded call turns. The override changes only the calls-enabled agent's LLM selection, while STT, TTS, prompts, history, memory, tools, and delegated-agent model selection remain unchanged.
models:
chat:
provider: anthropic
id: claude-opus-5
call_fast:
provider: anthropic
id: claude-haiku-4-5
agents:
assistant:
display_name: Assistant
model: chat
calls:
enabled: true
profiles:
fast-cascaded:
backend: cascaded
model: call_fast
stt:
provider: openai
model: gpt-transcribe
credentials_service: openai-voice
tts:
provider: openai
model: gpt-4o-mini-tts
credentials_service: openai-voice
extra_kwargs:
voice: ash
agents:
assistant: fast-cascaded
The configured call_fast alias is validated when MindRoom loads the configuration.
An unknown alias fails configuration loading instead of falling back to the normal agent model.
Cascaded cloud example
This example omits the cascaded model override and uses OpenAI speech services while the agent keeps its normally resolved model.
The agent can therefore use Anthropic, Gemini, OpenAI, or any other MindRoom model provider without changing the call configuration.
calls:
enabled: true
profiles:
openai-cascaded:
backend: cascaded
stt:
provider: openai
model: gpt-transcribe
credentials_service: openai-voice
extra_kwargs:
language: en
tts:
provider: openai
model: gpt-4o-mini-tts
credentials_service: openai-voice
extra_kwargs:
voice: ash
agents:
assistant: openai-cascaded
Each speech component has its own provider, model, credentials_service, api_key, host, and extra_kwargs.
The two speech legs may select the same named credential or different ones.
The current OpenAI model catalog documents gpt-transcribe for transcription and gpt-4o-mini-tts for text-to-speech.
Completely local example
This configuration sends STT, LLM, and TTS requests only to separately managed localhost services and needs no real cloud API key.
The STT service must expose an OpenAI-compatible /v1/audio/transcriptions endpoint.
The TTS service must expose an OpenAI-compatible /v1/audio/speech endpoint, such as a narrow local Kokoro server using a configured model alias.
models:
default:
provider: openai
id: local-chat-model
extra_kwargs:
api_key: sk-no-key-required
base_url: http://127.0.0.1:9292/v1
memory: none
agents:
assistant:
display_name: Assistant
role: Help the caller.
rooms: [lobby]
calls:
enabled: true
profiles:
local-cascaded:
backend: cascaded
stt:
provider: openai_compatible
model: whisper-large-v3
host: http://127.0.0.1:9000
extra_kwargs:
language: en
tts:
provider: openai_compatible
model: tts-1
host: http://127.0.0.1:9001
extra_kwargs:
voice: ash
agents:
assistant: local-cascaded
host accepts either the service root or its /v1 base URL.
MindRoom supplies a non-secret placeholder key when an OpenAI-compatible speech endpoint has no configured key.
LiveKit's local Silero VAD controls turn boundaries and barge-in, and preemptive agent generation stays disabled so an interrupted or speculative turn cannot execute tools twice.
Server requirements
Your Matrix deployment needs the standard Element Call backend:
- A LiveKit SFU reachable by call participants.
- The MatrixRTC authorization service (
lk-jwt-service) that exchanges Matrix OpenID tokens for LiveKit JWTs. - The Matrix server-name domain's
.well-known/matrix/clientmust advertise the service:
{
"org.matrix.msc4143.rtc_foci": [
{ "type": "livekit", "livekit_service_url": "https://rtc.example.org" }
]
}
Element's self-hosting guide covers the full setup, and matrix-docker-ansible-deploy enables all of it with matrix_rtc_enabled: true.
MindRoom joins only calls whose oldest membership advertises the locally configured or discovered MatrixRTC focus. It does not connect the server-hosted agent to participant-selected remote focuses; remaining participants may still inherit and advertise the trusted local focus after the original founder leaves.
The room's power levels must allow members to send org.matrix.msc3401.call.member state events (Element Call-capable clients set this up when they create rooms).
Homeserver notes
- Synapse supports the full MatrixRTC stack, including MSC4140 delayed events for automatic membership cleanup.
- Tuwunel works with Element Call but does not support delayed events yet (tuwunel#178), so memberships of crashed clients linger until their
expireswindow passes.
Encrypted rooms
Element Call encrypts call media with per-sender frame keys distributed over olm-encrypted to-device messages. MindRoom sends its own frame key this way, so participants can always hear the agent. Hearing the participants in an encrypted room requires mindroom-nio 0.27.0 or newer, which surfaces unknown decrypted to-device events and is required by MindRoom's dependency metadata (mindroom-nio#5). Calls in unencrypted rooms need none of this and work with plain SFU media.
Transcripts and memory
Every call writes a markdown transcript incrementally.
Shared file-memory agents keep it under calls/ in their canonical workspace, where their file tools can read it later; other shared agents keep it under <storage>/calls/<agent>/.
Requester-private file-memory agents keep transcripts in their requester-scoped workspace, while other requester-private agents keep them in their requester-scoped state archive.
When the call ends, file memory stores a relative transcript reference, Mem0 stores the transcript as recallable context, and disabled memory leaves only the transcript file.
Limitations
- Audio only: the agent neither publishes nor consumes video and screen shares.
- Managed agent calls support one distinct human Matrix user at a time, although that user may join from multiple devices.
- Cascaded speech services currently use OpenAI-compatible transcription and speech endpoints.
- Legacy 1:1
m.call.*calls (non-MatrixRTC) are not supported.