Skip to content

Voice Calls

MindRoom agents can join Element Call voice calls in their rooms and talk with you in real time. The realtime backend provides OpenAI realtime speech-to-speech. The live backend uses OpenAI Live for speech and delegates substantive requests to the normal MindRoom agent. The cascaded backend combines independently configured speech-to-text and text-to-speech services with the agent's normal MindRoom response path.

How it works

Matrix group calls (MatrixRTC, used by Element Call, Element X, and recent Cinny releases) do not send media over Matrix itself. Matrix only carries the signaling: participants publish org.matrix.msc3401.call.member state events, and media flows through a LiveKit SFU that is deployed next to the homeserver.

When a call starts in a room, the configured agent:

  1. Sees the call membership state event and re-reads the room state.
  2. Exchanges a Matrix OpenID token for a LiveKit JWT at the MatrixRTC authorization service (lk-jwt-service).
  3. Connects to the LiveKit SFU and publishes its own call membership state event, so it appears in the call roster.
  4. In encrypted rooms, distributes its media frame key over encrypted to-device messages and installs the other participants' keys, following the same per-sender key rotation policy as Element Call.
  5. Runs the selected voice backend until the managed session ends.

The voice agent is the same agent you chat with. Realtime carries the agent's rendered prompt and effective tools into OpenAI Realtime. Live receives the normal agent's full rendered system prompt, including its configured context files and caller-specific workspace, followed by voice delivery and delegation guidance. It can answer from that context directly and delegates requests needing tools, research, additional memory, or actions to the normal MindRoom agent. Tool schemas and execution remain in the delegated agent. The prompt is prepared before the first greeting, and the prepared agent is reused for delegated turns. The voice model relays the agent's results conversationally. Cascaded sends each finalized transcript through the normal MindRoom agent response path, preserving model resolution, the rendered system prompt, instructions, knowledge, skills, hooks, tools, requester identity, history storage, and tool execution behavior. A cascaded profile may explicitly override the resolved LLM while preserving all other agent behavior. MindRoom keeps a managed call active only while its devices belong to one Matrix user, and uses that user's ID as the real requester for the room-scoped tool runtime. Multiple devices belonging to that same user are supported. Tools requiring confirmation, user input, external execution, or tool_approval are hidden in all backends because a voice call has no approval UI.

The agent leaves the call and clears its membership state event when the caller leaves, a second distinct Matrix user joins, the voice session terminates, or the bot shuts down.

Configuration

calls:
  enabled: true
  profiles:
    openai-realtime:
      backend: realtime
      model: gpt-realtime-2.1
      credentials_service: openai-voice
      voice: marin
  agents:
    assistant: openai-realtime
  # livekit_service_url: https://rtc.example.org   # same-server .well-known override

Voice calls require the matrix_calls extra (pip install "mindroom[matrix_calls]" or uv sync --extra matrix_calls). calls.profiles defines reusable, complete voice pipelines. calls.agents maps each enabled agent to exactly one profile name. The model in a realtime or Live profile is the provider-specific OpenAI speech model ID. Live profiles require an explicit credentials_service and voice, and do not use separate STT or TTS services. The optional agent_model in a Live profile references a named entry from the top-level models mapping for delegated agent turns. Cascaded profiles require complete STT and TTS settings. OpenAI speech legs require either credentials_service or api_key; OpenAI-compatible speech legs require host and may omit credentials. The optional model in a cascaded profile references a named entry from the top-level models mapping. An explicit cascaded call model takes precedence over room and agent models for the calls-enabled agent. Omitting it keeps the existing room and agent model resolution. There are no inherited backend defaults or credential fallbacks. enabled and livekit_service_url remain global and cannot be overridden per agent. Set a dedicated credential service when voice and chat use different OpenAI credentials. A missing selected service does not fall back to another key. For example, these two agents use independent voice pipelines:

calls:
  enabled: true
  profiles:
    openai-realtime:
      backend: realtime
      model: gpt-realtime-2.1
      credentials_service: openai-voice
      voice: marin
    local-cascaded:
      backend: cascaded
      stt:
        provider: openai_compatible
        model: whisper-large-v3
        host: http://127.0.0.1:9000
      tts:
        provider: openai_compatible
        model: kokoro
        host: http://127.0.0.1:9001
        extra_kwargs:
          voice: af_heart
  agents:
    concierge: openai-realtime
    local_assistant: local-cascaded

MindRoom enforces at most one calls-enabled agent per room. Calls only join rooms configured for that agent and only while the sole caller passes the normal room and per-agent reply permissions. Calls-enabled agents also join calls in ad-hoc rooms they accepted through their normal authorized-invite policy. This lets Matrix clients create a private, temporary voice room and invite one agent without adding that room to config.yaml first. An explicit room assignment to one calls-enabled agent takes precedence over persisted invite records; invite-only ambiguity still fails closed. Calls-enabled agents advertise 📞 Voice calls in their Matrix presence status only when their MatrixRTC runtime is available, so clients can show the call action only where it will be answered. Requester-private agents use the sole authorized caller's verified Matrix user ID to resolve their workspace, memory, credentials, history, knowledge, and tool execution scope. The same caller scope applies in configured rooms and authorized ad-hoc invite rooms.

OpenAI Live with agent delegation

This example uses GPT-Live for speech and the configured chat model for the normal agent's delegated responses. The voice model uses its own named OpenAI credential service, independently of the agent's model provider and credentials.

models:
  chat:
    provider: anthropic
    id: claude-opus-5

agents:
  assistant:
    display_name: Assistant
    model: chat

calls:
  enabled: true
  profiles:
    openai-live:
      backend: live
      model: gpt-live-1
      credentials_service: openai-live
      voice: marin
      agent_model: chat
  agents:
    assistant: openai-live

agent_model is optional; omitting it keeps normal room and agent model resolution. When set, it overrides that resolution for the calls-enabled agent and must name an entry in models. An unknown alias fails configuration loading. The model field selects the OpenAI Live speech model and is independent of agent_model. The named credentials_service must resolve an API key; a missing service does not fall back to another OpenAI key. Live needs no separate STT or TTS credentials. Media reconnects reuse the delegate's history for the current call; once the caller leaves, the next call starts a new delegate session. Configuration reload rebuilds the call runtime when delegate model definitions or room model routing change.

Cascaded LLM model override

This example uses a large model for text conversations and a faster model for cascaded call turns. The override changes only the calls-enabled agent's LLM selection, while STT, TTS, prompts, history, memory, tools, and delegated-agent model selection remain unchanged.

models:
  chat:
    provider: anthropic
    id: claude-opus-5
  call_fast:
    provider: anthropic
    id: claude-haiku-4-5

agents:
  assistant:
    display_name: Assistant
    model: chat

calls:
  enabled: true
  profiles:
    fast-cascaded:
      backend: cascaded
      model: call_fast
      stt:
        provider: openai
        model: gpt-transcribe
        credentials_service: openai-voice
      tts:
        provider: openai
        model: gpt-4o-mini-tts
        credentials_service: openai-voice
        extra_kwargs:
          voice: ash
  agents:
    assistant: fast-cascaded

The configured call_fast alias is validated when MindRoom loads the configuration. An unknown alias fails configuration loading instead of falling back to the normal agent model.

Cascaded cloud example

This example omits the cascaded model override and uses OpenAI speech services while the agent keeps its normally resolved model. The agent can therefore use Anthropic, Gemini, OpenAI, or any other MindRoom model provider without changing the call configuration.

calls:
  enabled: true
  profiles:
    openai-cascaded:
      backend: cascaded
      stt:
        provider: openai
        model: gpt-transcribe
        credentials_service: openai-voice
        extra_kwargs:
          language: en
      tts:
        provider: openai
        model: gpt-4o-mini-tts
        credentials_service: openai-voice
        extra_kwargs:
          voice: ash
  agents:
    assistant: openai-cascaded

Each speech component has its own provider, model, credentials_service, api_key, host, and extra_kwargs. The two speech legs may select the same named credential or different ones. The current OpenAI model catalog documents gpt-transcribe for transcription and gpt-4o-mini-tts for text-to-speech.

Completely local example

This configuration sends STT, LLM, and TTS requests only to separately managed localhost services and needs no real cloud API key. The STT service must expose an OpenAI-compatible /v1/audio/transcriptions endpoint. The TTS service must expose an OpenAI-compatible /v1/audio/speech endpoint, such as a narrow local Kokoro server using a configured model alias.

models:
  default:
    provider: openai
    id: local-chat-model
    extra_kwargs:
      api_key: sk-no-key-required
      base_url: http://127.0.0.1:9292/v1

memory: none

agents:
  assistant:
    display_name: Assistant
    role: Help the caller.
    rooms: [lobby]

calls:
  enabled: true
  profiles:
    local-cascaded:
      backend: cascaded
      stt:
        provider: openai_compatible
        model: whisper-large-v3
        host: http://127.0.0.1:9000
        extra_kwargs:
          language: en
      tts:
        provider: openai_compatible
        model: tts-1
        host: http://127.0.0.1:9001
        extra_kwargs:
          voice: ash
  agents:
    assistant: local-cascaded

host accepts either the service root or its /v1 base URL. MindRoom supplies a non-secret placeholder key when an OpenAI-compatible speech endpoint has no configured key. LiveKit's local Silero VAD controls turn boundaries and barge-in, and preemptive agent generation stays disabled so an interrupted or speculative turn cannot execute tools twice.

Server requirements

Your Matrix deployment needs the standard Element Call backend:

  • A LiveKit SFU reachable by call participants.
  • The MatrixRTC authorization service (lk-jwt-service) that exchanges Matrix OpenID tokens for LiveKit JWTs.
  • The Matrix server-name domain's .well-known/matrix/client must advertise the service:
{
  "org.matrix.msc4143.rtc_foci": [
    { "type": "livekit", "livekit_service_url": "https://rtc.example.org" }
  ]
}

Element's self-hosting guide covers the full setup, and matrix-docker-ansible-deploy enables all of it with matrix_rtc_enabled: true.

MindRoom joins only calls whose oldest membership advertises the locally configured or discovered MatrixRTC focus. It does not connect the server-hosted agent to participant-selected remote focuses; remaining participants may still inherit and advertise the trusted local focus after the original founder leaves.

The room's power levels must allow members to send org.matrix.msc3401.call.member state events (Element Call-capable clients set this up when they create rooms).

Homeserver notes

  • Synapse supports the full MatrixRTC stack, including MSC4140 delayed events for automatic membership cleanup.
  • Tuwunel works with Element Call but does not support delayed events yet (tuwunel#178), so memberships of crashed clients linger until their expires window passes.

Encrypted rooms

Element Call encrypts call media with per-sender frame keys distributed over olm-encrypted to-device messages. MindRoom sends its own frame key this way, so participants can always hear the agent. Hearing the participants in an encrypted room requires mindroom-nio 0.27.0 or newer, which surfaces unknown decrypted to-device events and is required by MindRoom's dependency metadata (mindroom-nio#5). Calls in unencrypted rooms need none of this and work with plain SFU media.

Transcripts and memory

Every call writes a markdown transcript incrementally. Shared file-memory agents keep it under calls/ in their canonical workspace, where their file tools can read it later; other shared agents keep it under <storage>/calls/<agent>/. Requester-private file-memory agents keep transcripts in their requester-scoped workspace, while other requester-private agents keep them in their requester-scoped state archive. When the call ends, file memory stores a relative transcript reference, Mem0 stores the transcript as recallable context, and disabled memory leaves only the transcript file.

Limitations

  • Audio only: the agent neither publishes nor consumes video and screen shares.
  • Managed agent calls support one distinct human Matrix user at a time, although that user may join from multiple devices.
  • Cascaded speech services currently use OpenAI-compatible transcription and speech endpoints.
  • Legacy 1:1 m.call.* calls (non-MatrixRTC) are not supported.