Background Python Scripts
The script tool lets an agent start a Python program that keeps running after the chat turn ends.
Use it for watchers, polling loops, and other small automations that should wake the agent only when something changes.
The script can call a chosen subset of the agent's tools through the MindRoomTools SDK, with the same hooks, approval rules, worker routing, and audit trail as an ordinary tool call.
Treat script as trusted arbitrary code with authority comparable to python or shell.
allowed_tools limits only calls made through MindRoomTools; it does not restrict Python imports, filesystem access, operating-system calls, subprocesses, or direct network clients in the script.
Enable The Tool
Add script to the agent that will own the process, together with the toolkits the script may call.
This example lets a watcher read a URL and send a Matrix message to its own agent when the value changes.
models:
default:
provider: anthropic
id: claude-sonnet-5-5
agents:
watcher:
display_name: Watcher
role: Watch configured values and investigate meaningful changes.
model: default
rooms: [operations]
tools:
- script:
allowed_tools: [website, matrix_message]
max_concurrent_runs: 3
max_tool_calls_per_minute: 30
max_runtime_hours: 24
- website
- matrix_message
instructions:
- Use background scripts only for bounded, observable automation.
- Cancel watchers that are no longer needed.
defaults:
tools: []
Scripts also need a worker backend or explicit local execution; see Worker And Network Requirements.
| Option | Default | Meaning |
|---|---|---|
allowed_tools |
[] |
Toolkit names (not function names) the script may call through MindRoomTools |
max_concurrent_runs |
3 |
Maximum unfinished runs for one requester and agent, at least 1 |
max_tool_calls_per_minute |
30 |
Maximum new SDK calls per minute for one run, at least 1; automatic SDK retries and status polling do not count |
max_runtime_hours |
24 |
Maximum lifetime of one run in hours, positive and finite; the deadline holds even while the primary runtime is offline |
Each run keeps the limits and allowed_tools value in effect when it started.
Which tools a script can call and when it needs approval
- A non-empty
allowed_toolslist restricts the script to those toolkits, and their functions run without approval. - An empty list lets the script call every eligible toolkit on the agent, but each call waits for approval unless a
tool_approvalrule auto-approves it. - Unknown or ineligible toolkit names, including
*, makestart_scriptfail with an error. - The
browser,script,compact_context,delegate,dynamic_tools,dynamic_workflow,memory,self_config, andskill_managetoolkits are never available to scripts, even when the agent has them. - Listing
claude_agent,config_manager, orschedulerinallowed_toolsdoes not auto-approve their calls; only atool_approvalrule can. - Operator
tool_approvalrules are checked first, so a matchingrequire_approvalrule still pauses the call, and functions that declare their own confirmation still ask for it. - Only the original requester can approve or deny a script's call, and approval cards for a cancelled or finished run are settled without running the call.
Access is checked again on every call.
Removing a toolkit or function from the agent, or the requester losing access to the room or agent, takes effect immediately without restarting the script.
Adding toolkits later, including expanding allowed_tools, never gives a running script new unapproved access.
Control Functions
The agent receives five control functions.
start_script(
source: str | None = None,
path: str | None = None,
name: str | None = None,
resource_profile: Literal["small", "standard", "large"] | None = None,
)
get_script_resource_profiles()
get_script(run_id: str)
cancel_script(run_id: str, force: bool = False)
list_scripts(include_finished: bool = True)
start_scriptrequires exactly one ofsource(inline Python) orpath(relative to the agent workspace), limited to 128 KiB. Apathfile is copied at launch, so later edits do not change the running program.nameis an optional label shown in status and list results.resource_profileselects a Kubernetes CPU and memory profile; omitting it uses the configured default.get_script_resource_profilesshows the default and each profile's exact requests and limits, and start, status, and list results report the selected profile. An explicit profile fails on a backend that cannot enforce resource profiles. Operators define the profiles as described in Kubernetes deployment and the environment variable reference.get_scriptreturns the run state plus recent process output when available.list_scriptsreturns only runs owned by the current requester and agent.cancel_scriptimmediately revokes the script's tool access, requests graceful termination, force-kills after a short grace period, and reportscancelledonce the process has exited.force=Truekills the process without the graceful step.
Calling Agent Tools From A Script
The worker environment includes the mindroom.script_sdk client, which reads its connection settings from the environment the launcher provides.
Create one MindRoomTools instance and call tools by toolkit name, function name, and JSON-compatible keyword arguments.
from mindroom.script_sdk import MindRoomTools
tools = MindRoomTools()
result = tools.call("website", "read_url", url="https://example.org/status.txt")
print(result, flush=True)
MindRoomTools.call(toolkit_name, function_name, **arguments) blocks until the call finishes and returns a JSON-compatible value.
Arguments must have an unambiguous strict-JSON representation.
A result larger than about 64 KiB fails the call with kind result_too_large.
Calls within one run execute one at a time, in order.
The SDK retries transient failures automatically without repeating the side effect.
Failures raise MindRoomToolCallError, which exposes kind, retryable, and call_id so a script can log or stop predictably.
A denied approval raises a non-retryable error with kind approval_denied.
An indeterminate call means MindRoom accepted the call but cannot prove whether its side effect happened, so the script must not repeat it automatically.
Complete Watcher Example
This watcher polls a text endpoint and wakes the same agent once per observed value change. Replace the URL and agent name with values for your deployment.
from __future__ import annotations
import time
from mindroom.script_sdk import MindRoomTools
STATUS_URL = "https://example.org/controlled-status.txt"
AGENT_NAME = "watcher"
POLL_SECONDS = 15
def main() -> None:
tools = MindRoomTools()
previous = tools.call("website", "read_url", url=STATUS_URL)
while True:
time.sleep(POLL_SECONDS)
current = tools.call("website", "read_url", url=STATUS_URL)
if current == previous:
continue
previous = current
tools.call(
"matrix_message",
"matrix_message",
action="send",
message=f"The watched value changed; inspect {STATUS_URL} now.",
recipient=AGENT_NAME,
)
if __name__ == "__main__":
main()
The script runs in the room, thread, and requester context it was started from, so omitting room_id sends to that conversation, and sending to another room still requires normal Matrix authorization.
Setting recipient to the agent's own name starts a new turn for that agent; see Matrix Message for recipient rules.
Self-messaging requires a human requester, so start a self-triggering watcher from a turn a person started.
Make the watcher edge-triggered: update its observed value before sending, and never react to its own unchanged output.
Worker And Network Requirements
Run scripts on a dedicated Docker or Kubernetes worker backend as described in Sandbox Proxy.
The shared static_runner sandbox is not supported.
Workers should run the same MindRoom release as the primary; see Keeping worker versions matched.
Without a worker backend or explicit local execution, start_script fails with Background scripts require a worker or an explicitly disabled sandbox.
Workers must reach the primary's script gateway over the network.
Set MINDROOM_SCRIPT_GATEWAY_URL to the complete worker-reachable gateway base, including /api/script-gateway.
With Docker, MINDROOM_PUBLIC_URL can name the reachable MindRoom origin instead, and MindRoom appends /api/script-gateway.
Missing, malformed, credential-bearing, unresolvable, unspecified, or loopback gateway addresses are rejected, as are URLs with a query string or fragment.
export MINDROOM_WORKER_BACKEND=docker
export MINDROOM_DOCKER_WORKER_IMAGE=mindroom:dev
export MINDROOM_SANDBOX_PROXY_TOKEN=replace-with-a-long-random-token
export MINDROOM_SCRIPT_GATEWAY_URL=https://mindroom.example.org/api/script-gateway
Kubernetes
Kubernetes scripts stay disabled until the gateway runs on a listener that exposes nothing else, because workers that can reach the general API listener get more authority than the script gateway grants.
Until then, start_script fails with Kubernetes background scripts require a gateway-only listener; use Docker or explicit unsafe-local mode until that boundary is configured.
MINDROOM_SCRIPT_GATEWAY_PORTmakes the primary serve a second listener that answers only/api/script-gatewayroutes.MINDROOM_SCRIPT_GATEWAY_URLmust name that listener explicitly;MINDROOM_PUBLIC_URLis not used on Kubernetes.MINDROOM_SCRIPT_GATEWAY_ISOLATED=trueattests that workers cannot reach other primary API routes through that listener. The flag does not create network isolation, and the main API port is unchanged, so network policy must still keep workers off it.
The runtime Helm chart's scriptGateway.enabled value configures the listener, its <fullname>-script-gateway ClusterIP Service, the worker and control-plane NetworkPolicies, NO_PROXY, and all three variables, as described in the runtime chart README.
export MINDROOM_WORKER_BACKEND=kubernetes
export MINDROOM_SCRIPT_GATEWAY_PORT=8767
export MINDROOM_SCRIPT_GATEWAY_URL=http://mindroom-runtime-script-gateway.mindroom.svc.cluster.local:8767/api/script-gateway
export MINDROOM_SCRIPT_GATEWAY_ISOLATED=true
With Agent Vault enabled, the script's own direct network requests do not receive Agent Vault credential injection, while its MindRoomTools calls run through normal tool routing and keep the configured Agent Vault boundary.
What a script's worker can access
- Each run gets its own dedicated worker and filesystem root, even when other runs belong to the same requester and agent.
- The worker sees the requester-and-agent scoped workspace (see Filesystem isolation and agent data), and the script can read and modify files there; runs of the same requester and agent share that workspace, so they see each other's files in it.
- The script can read operator-provided worker environment values, including backend
extra_envand workspace environment overlays. - Global worker credentials are not copied into script workers;
MindRoomToolscalls get their normal authority through the gateway, and each call runs where that tool normally runs, not in the script's worker. - Worker mode isolates processes and files, not the network: direct network access follows the worker's Docker, host, and egress policy, and
allowed_toolsdoes not govern it. - Private agents must use
private.per: user_agent; other private scopes fail withBackground-script workers for private agents require private.per=user_agent.
Local execution
Setting MINDROOM_SANDBOX_EXECUTION_MODE to off, local, or disabled runs scripts in local execution mode instead of a worker.
Local execution is marked unsafe, runs under the primary host account, inherits the primary process environment, and provides no secret isolation.
It requires Linux and fails on other platforms with Background scripts require Linux process-group containment.
Use local execution only for trusted development scripts.
Lifecycle And Failure Semantics
| Run state | Meaning |
|---|---|
starting, running |
The run is launching or active |
exited |
The process returned exit code zero |
failed |
Launch failed or the process returned a nonzero exit code |
cancelled |
Cancellation was requested and the process has exited |
interrupted |
MindRoom stopped or lost the run for one of the reasons below |
Tool calls are pending, completed, failed, or indeterminate.
Cancelling a run cannot roll back a side effect a tool already started; such a call is reported indeterminate rather than cancelled or failed.
Restarts, upgrades, and reloads
Kubernetes runs keep running through primary restarts, compatible upgrades, and config reloads that do not affect their owner. An upgraded image applies only to new runs; existing runs stay on the image they started with. While MindRoom starts up or replaces its worker backend, new launches return an error and must be retried once MindRoom is ready, while SDK calls from running scripts are retried automatically. Keep the isolated gateway address stable across upgrades.
Docker and local runs are interrupted by every primary restart. Worker-backend replacement also interrupts Docker runs.
A run is also interrupted when:
- its worker or supervisor is lost, for example through node loss;
- its owning agent, its
scripttool, or the requester's authorization is removed; - a reload changes the owning agent's execution isolation or assigned knowledge base paths, or changes
defaults.worker_grantable_credentials; - any plugin reload happens, even for plugins the script does not use;
- a restart changes what the run depends on, such as the gateway URL, worker ownership or storage, the selected resource profile's resources, or the
MINDROOM_SCRIPT_GATEWAY_ISOLATEDattestation.
Changing an unused resource profile does not interrupt runs. MindRoom never relaunches an interrupted script automatically, so design watchers to checkpoint their observed state for a deliberate relaunch.
If the Docker daemon, worker storage root, or worker name prefix changed while MindRoom was offline, script launches stay blocked until the old runs are retired. Restore the previous settings, restart once so MindRoom can retire those runs, then apply the new settings.
Output and retention
get_script(run_id) keeps returning up to 64 KiB of process output after the run ends.
Finished runs, their tool-call receipts, and approval records are kept for 30 days by default.
Set MINDROOM_SCRIPT_RETENTION_SECONDS to change the window; see the environment variable reference.