Knowledge Bases
Knowledge bases give agents access to your own documents.
Point a knowledge base at a folder or Git repository and assign it to agents.
In semantic mode (the default), MindRoom indexes the files with an embedder into a vector database, and assigned agents get a search_knowledge_base tool.
In files mode, MindRoom skips embeddings and agents inspect the files directly with their file and shell tools.
Quick Start
Add a knowledge base and assign it to an agent:
knowledge_bases:
docs:
description: Product documentation, support notes, and internal operating procedures
mode: semantic
path: ./knowledge_docs
watch: false
chunk_size: 5000
chunk_overlap: 0
agents:
assistant:
display_name: Assistant
role: A helpful assistant with access to our docs
knowledge_bases: [docs]
Place files in ./knowledge_docs/, then trigger a reindex from the dashboard or API, or set watch: true so MindRoom picks up file changes automatically.
Knowledge base IDs are the keys under knowledge_bases.
An ID must be a non-empty single path component such as docs or company_docs; ., .., and names containing /, \, or line breaks are rejected.
Files-Only Knowledge Base
Use mode: files when agents should search, grep, list, and read the source files directly without paying for embeddings.
knowledge_bases:
source_docs:
description: Source documents the agent should inspect directly
mode: files
path: ${MINDROOM_STORAGE_PATH}/agents/assistant/workspace/source_docs
agents:
assistant:
display_name: Assistant
role: A helpful assistant with direct file access to source docs
memory_backend: file
tools: [file, shell]
knowledge_bases: [source_docs]
A files-mode base has no vector index, does not use the embedder, and does not provide search_knowledge_base.
The agent needs a workspace, which memory_backend: file creates at ${MINDROOM_STORAGE_PATH}/agents/<agent>/workspace (private agents use their private root).
When the base's path is inside that workspace and the folder exists, MindRoom exposes it as knowledge/<base_id> and tells the agent to use file tools on it.
Bases outside the workspace are not linked, and with the default file_access: workspace the agent's file tools cannot reach them.
Configuration
knowledge_bases:
my_docs:
description: Product documentation, support notes, and internal operating procedures
mode: semantic # "semantic" builds a vector search index; "files" skips embeddings
path: ./knowledge_docs/my_docs # Folder containing documents
watch: false # Direct external edits require reindex; dashboard/API changes still refresh
require_content_before_publish: false
chunk_size: 5000
chunk_overlap: 0
| Field | Type | Default | Description |
|---|---|---|---|
description |
string | "" |
What the base contains; shown to agents in the search_knowledge_base tool description and in files-mode instructions so they know when to use it |
mode |
semantic or files |
semantic |
semantic builds an embedding-backed search index; files skips embeddings and exposes the files to workspace-aware agents |
path |
string | ./knowledge_docs |
Folder path, relative to the config file directory or absolute |
watch |
bool | true |
Watch the local folder and refresh in the background when files change. When false, direct external edits require an explicit reindex; dashboard/API uploads and deletes still trigger a refresh |
require_content_before_publish |
bool | false |
Keep a new semantic index in the initializing state until at least one source file exists, instead of publishing an empty index |
chunk_size |
int | 5000 |
Maximum characters per chunk for plain-text and Markdown files (minimum 128); other formats such as JSON and PDF use their reader's own chunking; semantic mode only |
chunk_overlap |
int | 0 |
Overlap characters between adjacent plain-text and Markdown chunks; must be smaller than chunk_size; semantic mode only |
include_patterns |
list | [] |
Root-anchored glob patterns a file must match to be included |
exclude_patterns |
list | [] |
Root-anchored glob patterns removed after include filtering |
include_extensions |
list | null |
Exact set of extensions to index instead of the default text-like set (semantic mode only) |
extra_extensions |
list | [] |
Extensions to index in addition to the default set (semantic mode only) |
exclude_extensions |
list | [] |
Extensions to exclude, applied last (semantic mode only) |
skip_hidden |
bool | true |
Skip files and folders whose names start with ., such as temp files from editors that write in place. Git-backed bases use git.skip_hidden instead |
git |
object | null |
Sync the folder from a Git repository; see Git-Backed Knowledge Bases |
Set skip_hidden: false if the folder intentionally contains dot-prefixed files that should be indexed.
Use a smaller chunk_size when your embedding server has lower token or batch limits; chunks that are too large make indexing fail with embedder 500 errors.
Keeping Indexes Fresh
Chat never waits for indexing or Git sync. Agents search the last successfully published index, and a missing, stale, or failed base schedules a background refresh while the current request continues. A successful refresh replaces the published index; a failed refresh keeps the previous one and shows the error in the base's status.
A refresh reuses the stored vectors of unchanged files and only embeds files that were added or edited.
Changing chunk_size, chunk_overlap, the embedder, or any file filter re-embeds the whole base on the next refresh.
An explicit reindex from the dashboard or API rebuilds every vector, which is how you replace an index you no longer trust.
An interrupted or failed semantic build continues where it stopped on the next refresh instead of starting over.
Temporary embedding failures (timeouts, connection errors, HTTP 408, 429, and 5xx) are retried, while permanent failures such as authentication errors stop the refresh until you fix the cause.
Knowledge base configuration hot-reloads.
After a chunk_size or chunk_overlap change, agents keep searching the previous index until the rebuild succeeds.
After a change to path, mode, the embedder, the Git source, or any file filter, the base is not searchable until its new index is built.
Indexing Performance
| Environment variable | Default | Description |
|---|---|---|
MINDROOM_KNOWLEDGE_FILE_INDEX_CONCURRENCY |
4 |
Files indexed concurrently within one semantic refresh, from 1 through 128; invalid values are rejected at startup |
MINDROOM_KNOWLEDGE_REFRESH_CONCURRENCY |
1 |
Knowledge base refreshes that run at the same time; must be at least 1 |
MINDROOM_KNOWLEDGE_REFRESH_SUBPROCESS_TIMEOUT_SECONDS |
3600 |
Maximum seconds for one background refresh; must be a finite positive number. Raise it for very large corpora |
Local sentence_transformers embedding indexes one file at a time regardless of MINDROOM_KNOWLEDGE_FILE_INDEX_CONCURRENCY.
File Type Filtering
In semantic mode, MindRoom indexes a default set of text-like extensions covering Markdown, plain text, source code, and structured text formats such as .json, .yaml, and .csv.
knowledge_bases:
reports:
path: ./knowledge_docs/reports
extra_extensions: [".pdf", ".docx"] # Default text-like set plus PDF and Word
exclude_extensions: [".csv"] # Drop CSV files from the default set
extra_extensionsadds extensions on top of the default set.include_extensionsreplaces the default set entirely, andextra_extensionsstill adds on top of it.exclude_extensionsremoves extensions last and wins over both.
The allowed set depends only on this configuration, never on the embedder or installed packages.
Formats such as .pdf, .docx, or .pptx need their reader package (for example pypdf, python-docx, or python-pptx) installed in the MindRoom environment.
When a reader package is missing, the refresh indexes the other files, logs a warning naming the file and the missing package, and reports an indexing error instead of publishing an index that silently lacks those files.
Extension filtering does not apply in files mode.
Semantic indexing skips symlinked files and folders inside a base, so copy linked documents into the base instead.
Files larger than 64 MiB are left out of every knowledge base with a logged warning, and the next refresh removes any vectors they already had.
A .json file is split into one searchable document per top-level JSON value.
If the file is not valid JSON, its raw text is indexed like plain text and a warning names the file and the line and column of the parse error; fix that error to get per-value splitting back.
Multiple Knowledge Bases
Define several knowledge bases and assign each agent the ones it needs:
knowledge_bases:
engineering:
path: ./knowledge_docs/engineering
product:
path: ./knowledge_docs/product
legal:
path: ./knowledge_docs/legal
chunk_size: 1000
chunk_overlap: 100
agents:
developer:
display_name: Developer
role: Engineering assistant
knowledge_bases: [engineering]
pm:
display_name: Product Manager
role: Product planning assistant
knowledge_bases: [product, engineering]
compliance:
display_name: Compliance
role: Legal and compliance reviewer
knowledge_bases: [legal]
When an agent searches several semantic bases, results are interleaved so no single base dominates the top results.
Knowledge base paths must not be nested inside one another.
Several bases may use exactly the same path to expose different filtered views of one folder, as long as their Git settings (repo_url, branch, credentials_service, and lfs) match or none of them uses Git; see File Filtering with Patterns.
Private Agent Knowledge
Use agents.<name>.private.knowledge when each requester of a private agent should have their own knowledge index built from their private root.
knowledge_bases:
company_docs:
description: Shared company policies, project notes, and operating procedures
path: ./company_docs
watch: false
agents:
mind:
display_name: Mind
role: A persistent personal AI companion
model: sonnet
private:
per: user
root: mind_data
template_dir: ./mind_template
knowledge:
description: Requester-private notes, preferences, and working memory for this agent
path: memory
watch: false
knowledge_bases: [company_docs]
Each requester's private knowledge path becomes <their private root>/memory, and their index is never shared with another requester.
The fields are listed under Private Fields.
private.knowledge.path is required unless private.knowledge.enabled is false; it must be relative to the private root (. allowed) and cannot be absolute or escape with ...
Private knowledge has no background file watcher: with watch: true or git, it refreshes when the agent uses it, and with watch: false and no git, external edits are not picked up automatically.
An agent can combine private knowledge with shared top-level knowledge_bases, which stay the mechanism for documents shared across agents or users.
Private knowledge is available in normal agent conversations, not through the OpenAI-compatible /v1 API.
With private.knowledge.git, use a dedicated subtree such as kb_repo; MindRoom rejects . and any path that overlaps private file memory (memory/, MEMORY.md) or content copied from template_dir.
Git-Backed Knowledge Bases
A knowledge base can sync its folder from a Git repository.
MindRoom refreshes a shared Git-backed base every poll_interval_seconds, starting one interval after startup, and agents keep using the last published index while a refresh runs.
knowledge_bases:
pipefunc_docs:
path: ./knowledge_docs/pipefunc
chunk_size: 1200
chunk_overlap: 120
git:
repo_url: https://github.com/pipefunc/pipefunc
branch: main
poll_interval_seconds: 300
lfs: false
skip_hidden: true
include_patterns:
- "docs/**"
Git Configuration Fields
| Field | Type | Default | Description |
|---|---|---|---|
repo_url |
string | required | Repository URL to clone and fetch; see the accepted forms below |
branch |
string | main |
Branch to track |
poll_interval_seconds |
int | 300 |
Interval between background Git refreshes (minimum 5) |
credentials_service |
string | null |
Credential service for private repositories; see Private Repository Authentication |
lfs |
bool | false |
Download Git LFS files after each sync. Requires git-lfs on the machine running MindRoom; bundled container images include it |
sync_timeout_seconds |
int | 3600 |
Abort a single Git command after this many seconds (minimum 5) |
skip_hidden |
bool | true |
Skip files and folders whose names start with . |
include_patterns |
list | [] |
Root-anchored glob patterns to include |
exclude_patterns |
list | [] |
Root-anchored glob patterns to exclude |
Accepted repo_url forms
| Form | Example |
|---|---|
| URL with a host | https://github.com/org/repo.git, ssh://git@host/org/repo.git |
| scp-style SSH | git@github.com:org/repo.git, github.com:org/repo.git |
| Absolute local path | /srv/repos/repo.git |
file: URL |
file:///srv/repos/repo.git |
Network forms must have a resolvable host.
MindRoom refuses, with an error naming the reason, relative and home-relative paths (./repo.git, ../repo.git, ~/repo.git, repo.git), and URLs with an empty host (https:///org/repo.git).
Do not embed credentials in repo_url.
A password in a well-formed URL is kept out of the checkout's Git configuration, but it stays in your config file and is passed to Git on every command.
A credential MindRoom cannot reliably separate from the URL, such as oauth2:TOKEN@gitlab.com:org/repo.git or x-access-token:TOKEN@github.com/org/repo.git, makes the sync fail.
Use credentials_service instead; those credentials reach Git only for the duration of each command and are never written to disk.
Sync Behavior
- An explicit reindex or sync from the dashboard or API syncs Git first, then rebuilds the semantic index or, in files mode, publishes the new file list.
- The checkout is owned by MindRoom: when the branch moves, sync discards local edits to tracked files and restores deleted ones, so edit the source repository and sync instead.
- Dashboard and API file uploads and deletes are rejected for Git-backed bases.
- When
lfs: true, MindRoom downloads all LFS files after each sync, even ones the indexing filters exclude. - LFS downloads use the endpoint derived from
repo_url; an LFS URL set in the repository's.lfsconfigis ignored.
File Filtering with Patterns
Patterns are matched from the repository or folder root.
* matches one path segment and ** matches zero or more segments.
knowledge_bases:
project_docs:
path: ./knowledge_docs/project
git:
repo_url: https://github.com/org/project
include_patterns:
- "docs/**" # All files under docs/
- "README.md" # Root README only
- "content/posts/*/index.md" # Specific nested files
exclude_patterns:
- "docs/internal/**" # Exclude internal docs
- If
include_patternsis empty, all non-hidden files are eligible. - If
include_patternsis set, a file must match at least one pattern. exclude_patternsare applied last and remove matching files.- On a Git-backed base, a file must pass both the top-level patterns and the
gitpatterns.
Pointing several bases at the same path is the preferred way to expose separate views of a large repository without cloning it more than once:
knowledge_bases:
project_docs:
path: ./knowledge_docs/project
git:
repo_url: https://github.com/org/project
branch: main
include_patterns: ["docs/**"]
project_source:
path: ./knowledge_docs/project
git:
repo_url: https://github.com/org/project
branch: main
include_patterns: ["src/**"]
Checkout Layout
The knowledge folder holds only the checked-out files.
The Git data lives at <storage>/knowledge_git/<folder>_<path digest>, and bases that share a folder share it.
A .git written into the knowledge folder is ignored and never indexed.
The folder must be empty when MindRoom first clones into it.
Changing a base's path starts a new clone in the new folder; delete the old directory under <storage>/knowledge_git/ to reclaim its space.
Deleting only the knowledge folder restores its files from the Git data on the next sync; delete both to clone afresh.
Knowledge Git commands ignore system and global Git configuration, repository hooks, and credential helpers, and they do not receive MindRoom's secrets.
Configure proxies, CA bundles, and SSH through environment variables or ~/.ssh/config, and credentials through credentials_service.
Private Repository Authentication
For private HTTPS repositories, store credentials and reference them in the config.
Step 1: Store credentials through the dashboard Credentials tab or the API:
curl -X POST http://localhost:8765/api/credentials/github_private \
-H "Content-Type: application/json" \
-d '{"credentials":{"username":"x-access-token","token":"ghp_your_token_here"}}'
Step 2: Reference the service name in the knowledge base config:
knowledge_bases:
private_docs:
path: ./knowledge_docs/private
git:
repo_url: https://github.com/org/private-repo
credentials_service: github_private
| Fields | Notes |
|---|---|
username + token |
Standard GitHub or GitLab access token |
username + password |
Basic HTTP auth; the password wins if a token is also set |
token or api_key alone |
Authenticates as x-access-token |
For unattended GitHub access, a GitHub App installation can replace a personal access token.
Store this object under the service named by credentials_service:
{
"auth_type": "github_app",
"app_id": 12345,
"installation_id": 67890,
"private_key_file": "/var/run/secrets/github-app/private-key.pem"
}
private_key_file must be an absolute path, ideally a read-only secret mount; do not put the PEM contents in the credential object.
repo_url must use the canonical https://github.com/<owner>/<repository> form.
MindRoom requests short-lived installation tokens limited to that repository with read-only Contents permission, so the App needs Contents read access, and the tokens are never written to disk or logs.
Open Knowledge Format Bundles
An Open Knowledge Format (OKF) bundle is a folder of markdown files whose YAML frontmatter records who wrote each concept, who verified it, and whether it is current.
Use mode: files for OKF bundles, because semantic search cannot filter on frontmatter.
Agents with the bundled open-knowledge-format skill report trust and freshness when they answer, and when they edit, they record themselves as the author and add a verified entry only when a person confirms the content.
Reading a Published Bundle
Mirror the bundle's repository inside the agent workspace and name the bundle root in the description:
knowledge_bases:
acme_okf:
description: Acme Retail finance knowledge, a read-only Open Knowledge Format (OKF) bundle mirrored from Git. Start at knowledge/acme_okf/bundles/acme_retail/index.md.
mode: files
path: ${MINDROOM_STORAGE_PATH}/agents/analyst/workspace/okf/acme
git:
repo_url: https://github.com/GoogleCloudPlatform/open-knowledge-format
agents:
analyst:
display_name: Analyst
role: Answers questions from the team's knowledge bundles
memory_backend: file
tools: [file]
knowledge_bases: [acme_okf]
skills: [open-knowledge-format]
The mirror is a Git-backed base, so trigger a sync from the dashboard to clone it before the first poll.
Maintaining a Bundle with Agents
For a bundle that agents edit, create a local folder with a root index.md inside the agent workspace, then add it to the agent's knowledge_bases:
knowledge_bases:
ops_wiki:
description: The team's ops wiki, an Open Knowledge Format (OKF) bundle that the team and its agents maintain.
mode: files
path: ${MINDROOM_STORAGE_PATH}/agents/analyst/workspace/okf/ops_wiki
Embedder Configuration
Semantic knowledge bases use the embedder configured under memory.embedder; files-mode bases use none.
Semantic knowledge and file-memory indexes support the openai, ollama, and sentence_transformers providers, while the Mem0 memory backend also supports huggingface.
| Provider | Model example | Notes |
|---|---|---|
openai |
text-embedding-3-small |
Also works with keyless OpenAI-compatible local endpoints |
ollama |
nomic-embed-text |
Self-hosted; set host or OLLAMA_HOST |
sentence_transformers |
sentence-transformers/all-MiniLM-L6-v2 |
Runs locally in MindRoom; the optional extra installs on first use |
See Memory for the embedder fields, API key resolution, and how embedder failures are reported.
Storage
Semantic indexes are stored under <storage_path>/knowledge_db/<base_id>_<hash>/, where the hash comes from the resolved knowledge path.
For private agent knowledge, the requester's private root is part of that path, so each requester gets a separate index.
Files-mode bases need no vector database.
The storage path defaults to mindroom_data/ next to config.yaml and can be changed with MINDROOM_STORAGE_PATH.
Search Limits
A search times out after 30 seconds, and at most four searches run at once; a search that finds no free slot can fail with Knowledge reader is busy; try again shortly.
Dashboard and API
Manage knowledge bases, files, reindexing, and Git sync from the dashboard Knowledge tab or the API; see Dashboard for the tab and the knowledge API endpoints.