Memory System
Memory lets agents remember facts, preferences, and context across conversations.
This page covers choosing a memory backend, configuring embeddings and extraction, file-based memory, the memory tool, and Agno Learning.
| Backend | What it does | Use it when |
|---|---|---|
mem0 (default) |
Vector memory: relevant memories are retrieved before each reply and new ones are extracted automatically after each turn | You want hands-off semantic memory |
file |
Markdown files (MEMORY.md plus dated notes in memory/) that are the source of truth |
You want memory you can read and edit, or an OpenClaw-style workspace |
none |
No built-in memory | The agent should keep no memories (Agno Learning stays on unless disabled) |
Choosing a Backend
Set the global backend with memory.backend (mem0, file, or none; default mem0), and override it for one agent with agents.<name>.memory_backend.
memory: none is shorthand for memory: {backend: none}.
memory:
backend: mem0
agents:
coder:
display_name: Coder
role: Write and review code
memory_backend: file
scratch:
display_name: Scratch
role: Stateless scratchpad agent
memory_backend: none
With none, the agent gets no memories in its prompt, stores none after its turns, and does not receive the memory tool.
none does not turn off Agno Learning; set learning: false as well to disable all built-in memory and learning.
Conversation history is separate and stays in the session; see History & Compaction.
OpenClaw-compatible agents use this same backend selection.
Memory Scopes
| Scope | ID | Contents |
|---|---|---|
| Agent | agent_<name> |
The agent's preferences and durable user context |
| Team | team_<agent1>+<agent2>+... (member names sorted) |
Shared memory from team conversations |
An agent's memory reads cover its own scope and the scopes of the configured teams it belongs to, searched through the agent's own backend.
A team reads its own team scope; set memory.team_reads_member_memory: true (default false) to let it also read its members' agent scopes.
Team memory picks one backend from the members' effective backends:
- All members use
file: the team uses file memory. - All members use
none: team memory is disabled. - Any other combination, including
fileplusnone, uses Mem0 with the configuredmemory.embedderandmemory.llm.
A member's none setting disables only its own memory, so team conversations it takes part in can still be remembered in the team scope.
A file member of a team that uses Mem0 does not see that team's memory in its own conversations.
Backend: mem0
Before each reply, Mem0 retrieves memories relevant to the message and adds them to the prompt. After each turn, the memory LLM extracts durable facts into a Chroma vector store under the agent's storage directory.
Embedder
memory.embedder sets the embedding model used by Mem0, semantic file-memory search, and knowledge bases.
Mem0 supports the openai (default), ollama, huggingface, and sentence_transformers providers; semantic file memory and knowledge bases support the narrower set.
| Field | Default | Description |
|---|---|---|
provider |
openai |
Embedding provider |
config.model |
text-embedding-3-small |
Embedding model name |
config.host |
null |
Base URL for OpenAI-compatible endpoints, or the Ollama host; for ollama, a configured OLLAMA_HOST takes precedence, and the fallback is http://localhost:11434 |
config.dimensions |
null |
Embedding dimension override (>= 1) for OpenAI-compatible providers |
config.api_key |
null |
Explicit API key for the openai provider |
config.credentials_service |
null |
Credential service that supplies the openai provider's key |
memory:
backend: mem0
embedder:
provider: openai
config:
model: text-embedding-3-small
credentials_service: embedder # Optional strict credential binding
dimensions: null # Optional, for example 256
Fully local embedders:
memory:
embedder:
provider: sentence_transformers
config:
model: sentence-transformers/all-MiniLM-L6-v2
MindRoom installs the optional sentence_transformers extra the first time that provider is used.
Changing the embedder provider, model, host, or dimensions switches Mem0 to a separate collection, so memories stored with the previous embedder are no longer retrieved.
The openai embedder takes its API key from the first of these that is set:
memory.embedder.config.api_key.- The service named by
memory.embedder.config.credentials_service; when it has no key, MindRoom does not fall back to another key. - When no service is named, the dedicated
embeddercredential (seeded fromEMBEDDER_API_KEYorEMBEDDER_API_KEY_FILE), then the sharedopenaiprovider key.
Use credentials_service when the embedding endpoint needs a different key than chat, voice, or other OpenAI-compatible models.
With no key at all, keyless local OpenAI-compatible endpoints still work, while a keyed endpoint fails with an authentication error.
When the embedder rejects its credential, semantic memory search reports the cause in memory tool output and in the agent's prompt instead of returning empty results, and /api/health includes an embedder block until a request succeeds again.
Memory LLM
memory.llm sets the model Mem0 uses to extract memories.
Supported providers are ollama, openai, openrouter, and anthropic.
When memory.llm is omitted, MindRoom uses gemma4 on the Ollama server set by OLLAMA_HOST (default http://localhost:11434).
With provider: ollama, config.host points extraction at a different Ollama server, but a configured OLLAMA_HOST takes precedence; other providers ignore host.
With provider: openai, use gpt-6-luna; MindRoom drops the temperature and top_p values Mem0 sends, which GPT-6 models reject.
OpenRouter-only example, where both the memory LLM and the embeddings use the openrouter key:
memory:
embedder:
provider: openai
config:
model: openai/text-embedding-3-small
host: https://openrouter.ai/api/v1
credentials_service: openrouter
dimensions: 1536
llm:
provider: openrouter
config:
model: openai/gpt-6-luna
The LLM uses memory.llm.config.api_key when set and not blank, and otherwise the provider's credential (for example openrouter, or a dashboard credential stored as OPENROUTER_API_KEY).
When OPENROUTER_API_KEY is set in the process environment, the openai and openrouter memory LLMs send that key to OpenRouter, overriding memory.llm.config.api_key.
Backend: file
File memory keeps MEMORY.md and dated notes under memory/ in the agent's workspace, and those files are the source of truth.
memory:
backend: file
file:
max_entrypoint_lines: 200
max_entrypoint_tokens: 50000
search:
mode: keyword
include:
- memory/**/*.md
include_entrypoint: false
auto_flush:
enabled: true
For a single agent, the file backend writes memory only when the agent uses the memory tool or edits the files, or when auto-flush is enabled; it does not extract memories after every turn like mem0.
A team whose members all use file appends each message it answers to the team's MEMORY.md, even without auto-flush.
Before each reply, MindRoom inlines the start of MEMORY.md into the prompt under a header with its path, so the agent does not need to re-read it, and adds memories that match the message.
memory.file.max_entrypoint_lines (default 200, minimum 1) caps the inlined lines, and memory.file.max_entrypoint_tokens (default 50000, minimum 1) caps their estimated size at characters / 4.
Only whole lines are inlined, so a line that would cross the token cap is withheld with everything after it.
When either cap truncates, a marker reports the included and total line counts, both caps, and the path to read for the rest.
memory.file.path (default unset, relative to the config directory) is a fallback root for team file memory and never moves agent file memory out of the workspace.
File layout
Agent file memory lives in the agent workspace, agents/<agent>/workspace/ in the storage directory:
MEMORY.md, whereadd_memorywritesmemory/YYYY-MM-DD.md, where auto-flush writes
Team file memory is mirrored under each member's storage directory:
agents/<agent>/memory_files/team_<sorted_members>/MEMORY.mdagents/<agent>/memory_files/team_<sorted_members>/memory/YYYY-MM-DD.md
Memory files must be regular files; symlinks, FIFOs, and device nodes are ignored when reading and rejected when writing. Each file is read up to 1 MiB, cut at its last complete line, and updating a truncated or non-UTF-8 file fails so the rest of it is preserved.
Searching file memory
memory.search controls how search_memories and per-turn retrieval search file memory, and agents.<name>.memory_search overrides it per agent, with omitted fields inherited.
| Field | Default | Description |
|---|---|---|
mode |
keyword |
keyword or semantic |
include |
[memory/**/*.md] |
Glob patterns relative to the file-memory root that semantic search indexes; must be non-empty |
include_entrypoint |
false |
Also index MEMORY.md for semantic search |
In keyword mode, search_memories matches structured entries with IDs in MEMORY.md, plus entries and plain-text lines in memory/**/*.md.
Ordinary prose in MEMORY.md reaches the agent through the prompt preload, not keyword search; to reach lines beyond the preload cap, read the file directly or use semantic search with include_entrypoint: true.
In semantic mode, MindRoom builds a vector index of the include files with memory.embedder on first use and stores it under <storage-root>/knowledge_db/ (see Knowledge storage).
Until the index is ready, or when embeddings fail, search falls back to keyword results.
Writes through the memory tool and auto-flush refresh the index, but a ready index does not notice direct file edits until a later such write refreshes it.
Semantic mode covers the agent's own memory; team file memory is always keyword searched.
agents:
coder:
display_name: Coder
role: Write and review code
memory_backend: file
memory_search:
mode: semantic
Once its semantic index is ready, an agent's file memory also appears as a read-only source in search_knowledge_base.
That source is semantic-only, without keyword fallback, team memory, or memory IDs, so use search_memories when you need those.
Per-requester file memory
To give each requester their own file memory from one agent definition, combine memory_backend: file with private; the requester's private root then becomes the file-memory root.
File Auto-Flush Worker
Auto-flush periodically turns recent conversation in file-backed agents into durable memories.
It applies only to agents whose effective backend is file.
- Each turn marks its session as having unsaved conversation.
- A background worker picks eligible sessions in bounded batches.
- The agent's own model reads the new messages, plus existing memories to avoid duplicates, and writes durable facts.
- When the model answers with the
no_reply_token, nothing is written. - Results are appended to
memory/YYYY-MM-DD.md.
memory:
backend: file
auto_flush:
enabled: true
flush_interval_seconds: 1800
idle_seconds: 120
max_dirty_age_seconds: 600
batch:
max_sessions_per_cycle: 10
extractor:
max_messages_per_flush: 20
| Field | Default | Description |
|---|---|---|
enabled |
false |
Turn on auto-flush |
flush_interval_seconds |
1800 (min 5) |
Time between worker cycles |
idle_seconds |
120 (min 0) |
Session idle time before it becomes eligible |
max_dirty_age_seconds |
600 (min 1) |
Make a session eligible after this long with unsaved conversation, even if still active |
stale_ttl_seconds |
86400 (min 60) |
Forget pending sessions older than this |
max_cross_session_reprioritize |
5 (min 0) |
Other pending sessions of the same agent moved up the queue when a new message arrives |
retry_cooldown_seconds |
30 (min 1) |
Wait before retrying a failed extraction |
max_retry_cooldown_seconds |
300 (min 1) |
Upper bound for the growing retry wait |
batch.max_sessions_per_cycle |
10 (min 1) |
Sessions processed per cycle |
batch.max_sessions_per_agent_per_cycle |
3 (min 1) |
Sessions per agent processed per cycle |
extractor.no_reply_token |
NO_REPLY |
Reply meaning nothing should be saved |
extractor.max_messages_per_flush |
20 (min 1) |
Messages considered per extraction |
extractor.max_chars_per_flush |
12000 (min 1) |
Message characters considered per extraction |
extractor.max_extraction_seconds |
30 (min 1) |
Timeout for one extraction, retried in a later cycle |
extractor.include_memory_context.memory_snippets |
5 (min 0) |
Existing memories shown to the model to avoid duplicates |
extractor.include_memory_context.snippet_max_chars |
400 (min 1) |
Characters per existing memory shown |
Prompt Curation
Agents append to MEMORY.md and their context files more often than they condense them, so the prompt they send on every turn keeps growing.
Enable the prompt_curation automation to have MindRoom check their size daily and, once they pass a trigger, ask the agent in a visible thread to condense them gradually and move detail into searchable memory/ files, then measure the result and ask the agent to re-check a cut outside the bounds.
[memory]
The memory tool lets an agent deliberately remember, look up, correct, or forget something, alongside automatic memory.
It has no tool-specific configuration; it always uses the agent's effective backend and the memory.search settings, and agents with memory_backend: none do not receive it.
add_memory("The user prefers terse release notes.")
search_memories("release notes", limit=3) # default limit 5
list_memories(limit=20) # default limit 50
get_memory("abc123")
update_memory("abc123", "The user prefers terse release notes with dates.")
delete_memory("abc123")
The tool reaches every agent and team scope visible to the agent, as described in Memory Scopes.
Search and list results include an ID: Mem0 memory IDs and file-backed file:<path>:<line> IDs work with get_memory, update_memory, and delete_memory.
Semantic file-memory matches use read-only semantic:<source_file>:<rank> locators that these functions do not accept.
Result metadata includes search_mode (keyword or semantic), showing which search produced each result.
Failures are returned as readable error messages rather than raised into the conversation.
UI Configuration
The Dashboard Memory page edits the memory section: backend, team_reads_member_memory, embedder provider, model, credential service, and host, file settings (path, max_entrypoint_lines, max_entrypoint_tokens), search settings, and all auto-flush settings.
The remaining fields, such as llm, are under More settings.
Save from the Memory page to write the changes to config.yaml.
Agno Learning
Agno Learning builds a persistent profile of user preferences from conversations, separately from the memory backends above.
It is on by default for every agent, and memory_backend: none does not disable it.
defaults:
learning: true # default
learning_mode: always # learn after every turn; agentic lets the agent decide via a tool
agents:
assistant:
learning: false # disable for one agent
research:
learning_mode: agentic
See Agent fields for how learning and learning_mode inherit from defaults.
Disabled agents create and update no learning data.
Learning data is stored in learning/<agent>.db under the agent's state root (agents/<agent>/ in the storage directory for shared agents), so it survives container restarts when the storage directory is mounted.