Web Scraping & Browser
This page covers the built-in tools that read web pages, extract article text, crawl sites, call hosted scraping APIs, and drive a browser. Use it to pick a tool, configure it, and understand why a URL was refused.
Choosing a Tool
| Tool | Use it for | Credentials |
|---|---|---|
website |
Reading one page quickly | None |
trafilatura |
Local text, metadata, HTML-to-text, batch extraction, and small crawls | None |
newspaper |
News articles and blog posts with title, authors, and publish date | None |
crawl4ai |
Local browser-rendered reads of one or a few URLs, optionally focused on a query | None |
jina |
Jina Reader page reads and optional search | Optional JINA_API_KEY |
firecrawl |
Hosted scrape, crawl, map, and search | FIRECRAWL_API_KEY |
spider |
Spider Cloud search, scrape, and crawl | SPIDER_API_KEY |
scrapegraph |
Prompt-driven structured extraction | SGAI_API_KEY |
apify |
Running an Apify Actor as a tool | APIFY_API_TOKEN |
brightdata |
Markdown scraping, screenshots, SERP queries, and data feeds | BRIGHT_DATA_API_KEY |
oxylabs |
Google SERP, Amazon product data, and generic scraping | OXYLABS_USERNAME and OXYLABS_PASSWORD |
agentql |
AgentQL query extraction in a local browser | AGENTQL_API_KEY |
browserbase |
Hosted remote browser sessions for navigation, screenshots, and page reads | BROWSERBASE_API_KEY and BROWSERBASE_PROJECT_ID |
browser |
Multi-step automation of MindRoom's browser or the user's own signed-in browser | None |
web_browser_tools |
Opening a URL in the host computer's browser for a person | None |
Each credential can be stored as the tool's api_key (or apify_api_token, username, and password) through the dashboard or credential store, or supplied as the environment variable shown.
Keep password fields out of inline YAML.
crawl4ai, agentql, browserbase, and browser need a working Playwright browser runtime.
For a visible, persistent browser that the user can watch or control in MindRoom Chat, see Worker Computer and Agent Chat UI Actions.
Network Access
Tools that fetch pages from MindRoom itself (website, trafilatura, newspaper, crawl4ai, agentql, and the browser host target) reach only public HTTP(S) addresses.
Private, loopback, link-local, multicast, reserved, and cloud-metadata targets are refused, and so are redirects, crawled links, and page resources that lead to them.
Unless a worker egress proxy is set, DNS answers that change after a URL is checked cannot redirect the connection to an internal service.
Only the browser tool can opt into private networks, with allow_private_networks.
In Computer mode, the browser tool also opens worker-local loopback previews such as http://localhost:5173 without that option; see Preview a local web app.
Egress Proxy for Browsers
crawl4ai, agentql, the browser host target, and browser_mcp in a Worker Computer send every connection through an operator egress proxy set with http_proxy, https_proxy, or all_proxy, using HTTP CONNECT.
In the primary process, http_proxy applies to port 80, https_proxy to every other port, and all_proxy to either when its own variable is unset.
The proxy must allow CONNECT to ports 80 and 443, because plain-HTTP pages are tunneled too; Squid's default http_access deny CONNECT !SSL_ports rule refuses plain HTTP.
In the primary process it must also allow CONNECT to IP addresses, so proxies that allow only hostnames are unsupported there.
Loopback destinations are always dialed directly, and no_proxy is honored only when allow_private_networks is enabled; it never makes a refused destination reachable.
A refused tunnel fails that connection without a direct fallback and logs browser_egress_proxy_refused_tunnel.
The primary process logs a warning and dials directly when the proxy is SOCKS, has credentials or a path in its URL, or is set through auto_proxy or socks_server.
In a worker, every set proxy variable must name the same HTTP(S) proxy, or the browser does not start; hostnames must also resolve in the worker's DNS.
That proxy receives hostnames and resolves them itself, so it must block private and metadata addresses and resist DNS rebinding.
WebRTC cannot send UDP from these browsers.
No-Key Scrapers
[website]
website exposes read_url(url) and returns JSON page documents from MindRoom's WebsiteReader variant, which drops search UI, navigation, headers, footers, sidebars, hidden content, and modals before choosing the page text.
Pages larger than 2 MiB fail, as do pages a server sends compressed.
| Option | Type | Default | Notes |
|---|---|---|---|
knowledge |
object |
null |
Programmatic Knowledge object only; it replaces read_url() with add_website_to_knowledge(url). |
[trafilatura]
trafilatura exposes extract_text(), extract_metadata_only(), crawl_website(), html_to_text(), and extract_batch().
extract_batch() returns one JSON payload listing successes and failures.
crawl_website() stops with an error when it reaches a refused target, while extract_content=True reports each refused page individually.
A server that sends a compressed body fails the download.
| Option | Type | Default | Notes |
|---|---|---|---|
output_format |
text |
txt |
txt, json, markdown, xml, csv, or html. |
include_comments |
boolean |
true |
Include comments. |
include_tables |
boolean |
true |
Include tables. |
include_images |
boolean |
false |
Include image information where supported. |
include_formatting |
boolean |
false |
Preserve formatting markers. |
include_links |
boolean |
false |
Preserve links. |
with_metadata |
boolean |
false |
Include metadata in extraction output. |
favor_precision |
boolean |
false |
Bias extraction toward precision. |
favor_recall |
boolean |
false |
Bias extraction toward recall. |
target_language |
text |
null |
ISO 639-1 language filter such as en. |
deduplicate |
boolean |
false |
Remove repeated segments. |
max_tree_size |
number |
null |
Parser tree-size limit. |
max_crawl_urls |
number |
10 |
Maximum URLs visited per crawl. |
max_known_urls |
number |
100000 |
Maximum discovered URLs tracked per crawl. |
enable_extract_text |
boolean |
true |
Enable extract_text(). |
enable_extract_metadata_only |
boolean |
true |
Enable extract_metadata_only(). |
enable_html_to_text |
boolean |
true |
Enable html_to_text(). |
enable_extract_batch |
boolean |
true |
Enable extract_batch(). |
enable_crawl_website |
boolean |
true |
Enable crawl_website(). |
all |
boolean |
false |
Enable every function. |
agents:
analyst:
tools:
- trafilatura:
output_format: markdown
with_metadata: true
include_links: true
[newspaper]
newspaper exposes read_article(url) and returns JSON with whichever of title, authors, text, and publish date were extracted.
It is tuned for article pages, not arbitrary sites, and returns no image fields.
Use newspaper in tools:, not newspaper4k.
| Option | Type | Default | Notes |
|---|---|---|---|
article_length |
number |
null |
Truncate article text to this many characters. |
enable_read_article |
boolean |
true |
Enable read_article(). |
all |
boolean |
false |
Enable every function. |
[crawl4ai]
crawl4ai exposes crawl(url, search_query=None), where url is one URL or a list, and returns readable text for each page from a local headless browser.
With search_query, the text is filtered toward that query; without one, use_pruning trims noisy content.
Results skip Crawl4AI's cache and are truncated to max_length.
| Option | Type | Default | Notes |
|---|---|---|---|
max_length |
number |
5000 |
Maximum returned characters. |
timeout |
number |
60 |
Crawl timeout in seconds. |
use_pruning |
boolean |
false |
Prune noisy content when no search_query is given. |
pruning_threshold |
number |
0.48 |
Pruning threshold. |
bm25_threshold |
number |
1.0 |
Query-filter threshold when search_query is given. |
headless |
boolean |
true |
Run the browser headless. |
wait_until |
text |
domcontentloaded |
Playwright wait condition before extraction. |
enable_crawl |
boolean |
true |
Enable crawl(). |
all |
boolean |
false |
Enable every function. |
[jina]
jina exposes read_url(url) through Jina Reader and, when enabled, search_query(query).
It works without a key for public reads; a key helps with paid plans and rate limits.
Returned content is truncated to max_content_length.
| Option | Type | Default | Notes |
|---|---|---|---|
api_key |
password |
null |
Optional; falls back to JINA_API_KEY. |
base_url |
url |
https://r.jina.ai/ |
Reader endpoint for read_url(). |
search_url |
url |
https://s.jina.ai/ |
Search endpoint for search_query(). |
max_content_length |
number |
10000 |
Maximum returned characters. |
timeout |
number |
null |
Jina timeout in seconds. |
search_query_content |
boolean |
true |
Include full page content in search results; false returns summaries only. |
enable_read_url |
boolean |
true |
Enable read_url(). |
enable_search_query |
boolean |
false |
Enable search_query(). |
all |
boolean |
false |
Enable every function. |
Hosted Scraping APIs
[firecrawl]
firecrawl exposes scrape_website(), crawl_website(), map_website(), and search_web().
It requires api_key or FIRECRAWL_API_KEY.
| Option | Type | Default | Notes |
|---|---|---|---|
api_key |
password |
null |
Falls back to FIRECRAWL_API_KEY. |
enable_scrape |
boolean |
true |
Enable scrape_website(). |
enable_crawl |
boolean |
false |
Enable crawl_website(). |
enable_mapping |
boolean |
false |
Enable map_website(). |
enable_search |
boolean |
false |
Enable search_web(). |
all |
boolean |
false |
Enable every function. |
formats |
string[] |
null |
Output formats for scrape, crawl, and search, such as markdown or html; check them against your plan. |
limit |
number |
10 |
Default result cap for crawl and search. |
poll_interval |
number |
30 |
Seconds between crawl status checks. |
api_url |
url |
https://api.firecrawl.dev |
Firecrawl API base URL. |
[spider]
spider exposes search_web(query, max_results=5), scrape(url), and crawl(url, limit=None), returning Markdown.
Search results do not include full page content.
It requires SPIDER_API_KEY in the environment even though the dashboard lists it as needing no setup, and it has no api_key field.
| Option | Type | Default | Notes |
|---|---|---|---|
max_results |
number |
null |
Default result count for search_web(). |
url |
url |
null |
Unused; scrape() and crawl() always take the URL as an argument. |
enable_search |
boolean |
true |
Enable search_web(). |
enable_scrape |
boolean |
true |
Enable scrape(). |
enable_crawl |
boolean |
true |
Enable crawl(). |
all |
boolean |
false |
Enable every function. |
[scrapegraph]
scrapegraph turns pages into structured answers from natural-language prompts.
smartscraper(url, prompt) extracts data from one page, markdownify() converts a page to Markdown, crawl() applies a prompt and JSON schema across a crawl, searchscraper() searches the web before extracting, and scrape() returns raw HTML.
It requires api_key or SGAI_API_KEY.
Keep at least one function enabled, or set all: true.
The upstream headers option is not supported in config.yaml.
| Option | Type | Default | Notes |
|---|---|---|---|
api_key |
password |
null |
Falls back to SGAI_API_KEY. |
enable_smartscraper |
boolean |
true |
Enable smartscraper(). |
enable_markdownify |
boolean |
false |
Enable markdownify(). |
enable_crawl |
boolean |
false |
Enable crawl(). |
enable_searchscraper |
boolean |
false |
Enable searchscraper(). |
enable_scrape |
boolean |
false |
Enable scrape(). |
render_heavy_js |
boolean |
false |
Render heavy JavaScript in scrape() only. |
crawl_poll_interval |
number |
3 |
Seconds between crawl status checks. |
crawl_max_wait |
number |
180 |
Maximum seconds to wait for a crawl. |
all |
boolean |
false |
Enable every function. |
[apify]
apify registers one tool function per configured Apify Actor, with parameters from the Actor's input schema, and returns the Actor's dataset items as JSON.
Without actors, it provides no functions.
Function names derive from the Actor ID, so check the agent's tool list for the exact name.
| Option | Type | Default | Notes |
|---|---|---|---|
apify_api_token |
password |
null |
Falls back to APIFY_API_TOKEN. |
actors |
text |
null |
One Actor ID, such as apify/rag-web-browser; a comma-separated list is treated as a single ID. |
[brightdata]
brightdata exposes scrape_as_markdown(), get_screenshot(), search_engine() for Google, Bing, and Yandex, and web_data_feed() for Bright Data's supported source types.
get_screenshot() returns the image directly to the model rather than a file path.
It requires api_key or BRIGHT_DATA_API_KEY.
| Option | Type | Default | Notes |
|---|---|---|---|
api_key |
password |
null |
Falls back to BRIGHT_DATA_API_KEY. |
enable_scrape_markdown |
boolean |
true |
Enable scrape_as_markdown(). |
enable_screenshot |
boolean |
true |
Enable get_screenshot(). |
enable_search_engine |
boolean |
true |
Enable search_engine(). |
enable_web_data_feed |
boolean |
true |
Enable web_data_feed(). |
all |
boolean |
false |
Enable every function. |
serp_zone |
text |
serp_api |
SERP zone; BRIGHT_DATA_SERP_ZONE overrides it. |
web_unlocker_zone |
text |
web_unlocker1 |
Web unlocker zone; BRIGHT_DATA_WEB_UNLOCKER_ZONE overrides it. |
verbose |
boolean |
false |
Log extra request detail. |
timeout |
number |
600 |
Timeout in seconds. |
[oxylabs]
oxylabs exposes search_google(), get_amazon_product(), search_amazon_products(), and scrape_website().
search_google() returns organic results with title, URL, description, and position.
Pass domain_code, such as com or de, to choose the regional Google or Amazon domain.
It requires both a username and password.
| Option | Type | Default | Notes |
|---|---|---|---|
username |
text |
null |
Falls back to OXYLABS_USERNAME. |
password |
password |
null |
Falls back to OXYLABS_PASSWORD. |
markdown |
boolean |
false |
Return scrape_website() content as Markdown instead of parsed HTML. |
Browser Tools
[agentql]
agentql exposes scrape_website(url), which extracts generic page text, and custom_scrape_website(url), which runs your agentql_query and returns the extracted values as JSON.
Setting agentql_query registers custom_scrape_website() even when enable_custom_scrape_website is false.
It opens only HTTP(S) URLs and launches a visible browser window, so it needs a GUI-capable runtime or virtual display.
It requires api_key or AGENTQL_API_KEY; AgentQL SDK settings and CLI credential files do not override that key.
| Option | Type | Default | Notes |
|---|---|---|---|
api_key |
password |
null |
Falls back to AGENTQL_API_KEY. |
enable_scrape_website |
boolean |
true |
Enable scrape_website(). |
enable_custom_scrape_website |
boolean |
false |
Enable custom_scrape_website(). |
all |
boolean |
false |
Enable every function. |
agentql_query |
text |
"" |
AgentQL query for custom_scrape_website(). |
[browserbase]
browserbase exposes navigate_to(), screenshot(), get_page_content(), and close_session() against a hosted Browserbase session that it creates automatically.
It still needs local Playwright to connect to the remote browser.
It requires an API key and project ID.
Use it when you need remote navigation, screenshots, and page reads without the full action set of browser.
| Option | Type | Default | Notes |
|---|---|---|---|
api_key |
password |
null |
Falls back to BROWSERBASE_API_KEY. |
project_id |
text |
null |
Falls back to BROWSERBASE_PROJECT_ID. |
base_url |
url |
null |
Browserbase API endpoint, not the site to visit; falls back to BROWSERBASE_BASE_URL. |
enable_navigate_to |
boolean |
true |
Enable navigate_to(). |
enable_screenshot |
boolean |
true |
Enable screenshot(). |
enable_get_page_content |
boolean |
true |
Enable get_page_content(). |
enable_close_session |
boolean |
true |
Enable close_session(). |
all |
boolean |
false |
Enable every function. |
parse_html |
boolean |
true |
Return visible text instead of raw HTML. |
max_content_length |
number |
100000 |
Maximum returned characters of page content. |
[browser]
browser exposes one function, browser_control(action=...), for multi-step browser sessions.
It can drive two targets:
target="host"(default) controls MindRoom's own Chromium on the MindRoom host, or in the agent's worker whenbrowseris listed inworker_tools.target="desktop"controls the user's own signed-in Chrome or Brave profile through the Matrix Desktop Bridge, which owns extension installation, the local control lease, and the trust model.
Call action="help" or action="actions" to list the actions and act request kinds.
| Action | Host | Desktop |
|---|---|---|
status, start, stop, profiles, tabs, open, snapshot, screenshot, console |
Yes | Yes |
focus, close, navigate, pdf, upload, dialog, act |
Yes | No |
help, actions |
Yes | Yes |
On the host target, snapshot() returns ai or aria format with element refs that later act() and screenshot() calls can use.
act() takes request.kind set to click, type, press, hover, drag, select, fill, resize, wait, evaluate, or close.
The desktop target returns the browser's native accessibility snapshot and rejects targetId and host-only options such as profile, snapshot format hints, inputRef, and timeoutMs.
Desktop-target calls always run in the primary process, so routing browser to a worker isolates only host-target calls.
To show the worker browser to the user, see Agent Chat UI Actions.
Screenshots and Files
Screenshots are shown to the model by default; host screenshots are also saved, and saveOnly=True saves without showing.
Large images may be resized before the model sees them.
Model-visible screenshots can be kept in the agent's session history.
On the desktop target, returnAttachment=True with action="screenshot" also returns an att_* handle, valid until the turn ends, that matrix_message can send.
Host screenshots, PDFs, and other artifacts go to output_dir.
In the primary process it defaults to browser/ in the agent's state root: <storage>/agents/<agent>/browser for a shared agent, or the requester's private-instance root for a private agent.
In a worker it defaults to browser/ in the worker workspace, and a custom output_dir must stay inside that workspace.
output_dir must not point at runtime state, and it does not affect the desktop target.
Which files upload may read follows the agent's file_access setting.
With the default workspace, upload accepts files in the artifact directory, files in the agent's workspace, and att_* IDs of attachments in the current conversation; use ./ for a workspace file whose name starts with att_.
Everything else is refused, including credentials, encryption keys, Matrix state, sessions, and other agents' workspaces; in a worker only the worker workspace is readable.
With unrestricted, upload accepts any file its process can read.
One browser holds at most 256 MiB of uploaded files until their tabs close, so a larger upload is refused until a tab is closed.
On the host target, open also takes one HTML file in paths instead of targetUrl, read under the same rules as upload, so an agent can check a page it wrote with screenshot, console (which also lists uncaught errors), and act.
The file is UTF-8 text of at most 16 MiB.
The page cannot load or navigate to other local files, and relative links in it do not resolve; its network requests follow the same checks as any page.
Profiles and Signed-In Sessions
Host-target profiles are named, with mindroom as the default; names starting with a dot are rejected.
In the primary process a profile lives at browser-profiles/<profile> under the agent's state root, so every requester of a shared agent shares its signed-in sessions, while different agents never share them.
This holds even with worker_scope: user or user_agent unless browser is in worker_tools, which gives each worker scope its own profiles.
Use a private agent when requesters need separate browser sessions.
Chromium is taken from BROWSER_EXECUTABLE_PATH, chromium, or google-chrome-stable.
Browser Configuration
| Option | Type | Default | Notes |
|---|---|---|---|
output_dir |
text |
null |
Host-target artifact directory; defaults to browser/ in the agent's state root, or in the worker workspace when routed to a worker. |
allow_private_networks |
boolean |
false |
Let the host target, and URLs passed to desktop-target open, reach trusted private and loopback addresses; metadata and link-local addresses stay blocked. Pages on the desktop target keep the profile's normal network access either way. |
default_target |
select |
host |
host or desktop. |
device_user_id |
text |
null |
Required for the desktop target: Matrix user the desktop bridge signs in as. |
device_id |
text |
null |
Required for the desktop target: device ID printed by mindroom desktop login. |
device_ed25519 |
text |
null |
Required for the desktop target: Ed25519 fingerprint of that device. |
timeout_seconds |
number |
90 |
Desktop-target timeout, from 1 to 120 seconds. |
agents:
browser_worker:
tools:
- browser:
default_target: desktop
device_user_id: "@my-laptop:example.org"
device_id: "ABCDEFGHIJ"
device_ed25519: "desktop-device-fingerprint"
browser_control(action="start", target="desktop")
browser_control(action="open", target="desktop", targetUrl="https://matrix.org/blog/")
browser_control(action="screenshot", target="desktop", fullPage=True, returnAttachment=True)
[web_browser_tools]
web_browser_tools exposes open_page(url, new_window=False), which opens an http or https URL in a tab or window of the host computer's browser for a person to use.
It refuses file: paths, other schemes, and scheme-less strings, and returns no page content, so it is not a scraper.
It only works on a host with a desktop browser.
| Option | Type | Default | Notes |
|---|---|---|---|
enable_open_page |
boolean |
true |
Enable open_page(). |
all |
boolean |
false |
Enable every function. |