Research Sources
Use these tools to query a specific source, such as ArXiv papers, Google Scholar citations, Wikipedia summaries, PubMed literature, or Hacker News, instead of general web search.
Tools On This Page
| Tool | Best for |
|---|---|
[arxiv] |
Searching ArXiv and reading the text of selected papers. |
[google_scholar] |
Cross-publisher publication search with citation counts and PDF links. |
[wikipedia] |
One encyclopedia summary per query. |
[pubmed] |
Medical and life-science literature. |
[hackernews] |
Top Hacker News stories and user profiles. |
Setup
None of these tools needs an API key or OAuth.
Add the tool to an agent's tools list; see Per-Agent Tool Configuration for inline options and include_tools/exclude_tools.
Missing Python dependencies install on first use; see Automatic Dependency Installation.
For toolkits with an all option, all: true enables every function regardless of the individual enable_* flags.
Use Web Search for broader web discovery or news search.
[arxiv]
arxiv searches ArXiv and can download papers to extract their page text.
| Option | Type | Default | Notes |
|---|---|---|---|
enable_search_arxiv |
boolean |
true |
Enable search_arxiv_and_return_articles(query, num_articles=10). |
enable_read_arxiv_papers |
boolean |
true |
Enable read_arxiv_papers(id_list, pages_to_read=None). |
all |
boolean |
false |
Enable both functions. |
download_dir |
text |
unset | Directory where downloaded PDFs are stored; when unset, PDFs go to an arxiv_pdfs directory inside the installed Agno package. |
Search returns JSON with each paper's title, ID, authors, categories, publish date, PDF URL, summary, and comment.
read_arxiv_papers() takes ArXiv IDs such as 2103.03404v1, not a search query, and returns the same metadata plus the text of each page; pages_to_read=None reads every page.
[google_scholar]
google_scholar provides search_google_scholar(query, max_results=None), which returns a JSON list with title, authors, year, venue, abstract, citation count, publication URL, and PDF URL when available.
| Option | Type | Default | Notes |
|---|---|---|---|
max_results |
number |
5 |
Result cap when the call does not pass max_results. |
Google Scholar has no official API, so results are scraped and citation counts, venues, and abstracts are best-effort.
Google Scholar rate-limits scrapers, and when it blocks requests the tool returns Google Scholar is currently blocking automated requests. Try again later.
Prefer arxiv or pubmed when they cover the topic, and use Google Scholar for cross-publisher coverage or citation counts.
[wikipedia]
wikipedia provides search_wikipedia(query), which returns the Wikipedia summary for the query as JSON.
An ambiguous query returns a list of candidate page titles instead of a summary, so retry with one of them or a more specific query.
| Option | Type | Default | Notes |
|---|---|---|---|
auto_suggest |
boolean |
true |
Let Wikipedia suggest or correct the title before lookup; set false for exact-title lookup. |
knowledge |
text |
unset | Has no effect in YAML configuration; leave unset. |
all |
boolean |
false |
Has no effect for this toolkit. |
[pubmed]
pubmed provides search_pubmed(query, max_results=None) through NCBI E-utilities and returns a JSON list of formatted text results.
By default each result has the title, publication year, and a summary truncated to about 200 characters.
With results_expanded: true, each result has the full abstract plus the first author, journal, publication type, DOI, PubMed URL, full-text URL when available, keywords, and MeSH terms.
| Option | Type | Default | Notes |
|---|---|---|---|
email |
text |
your_email@example.com |
Contact email sent to NCBI with each request; set a real address. |
max_results |
number |
unset | Result cap when the call does not pass max_results; 10 results when neither sets it. |
results_expanded |
boolean |
false |
Return the full abstract and the expanded metadata described above. |
enable_search_pubmed |
boolean |
true |
Enable search_pubmed(). |
all |
boolean |
false |
Enable all functions. |
timeout |
number |
30 |
Per-request HTTP timeout in seconds. |
agents:
clinician:
tools:
- pubmed:
email: research@example.com
max_results: 5
results_expanded: true
[hackernews]
hackernews reads the public Hacker News API.
get_top_hackernews_stories(num_stories=10) returns the current top story items, including title, URL, score, and author.
get_user_details(username) returns the user's karma, about text, and number of submitted items.
| Option | Type | Default | Notes |
|---|---|---|---|
enable_get_top_stories |
boolean |
true |
Enable get_top_hackernews_stories(). |
enable_get_user_details |
boolean |
true |
Enable get_user_details(). |
all |
boolean |
false |
Enable both functions. |
timeout |
number |
30 |
Per-request HTTP timeout in seconds. |
The tool returns story metadata only; pair it with Web Scraping & Browser to read the linked pages.