The Arxiv MCP Server enables AI assistants to search and access research papers from arXiv through the Model Context Protocol, providing tools to query papers with filters for date ranges and categories, download and read paper content, and list downloaded papers. It solves the problem of programmatically integrating arXiv's research repository with AI models by offering a standardized interface for paper discovery and access without requiring direct API management by the client.
claude mcp add arxiv --env ARXIV_STORAGE_PATH=YOUR_ARXIV_STORAGE_PATH -- uvx arxiv-mcp-server --storage-path '${ARXIV_STORAGE_PATH}'Run in your terminal. Replace YOUR_* placeholders with real values; add --scope user to install for every project.
Review the command, arguments, and environment values before installing — MCP servers run with your local permissions.
Verified live against the running server on Jun 10, 2026.
search_papersSearch for papers on arXiv with advanced filtering and query optimization. QUERY CONSTRUCTION GUIDELINES: - Use QUOTED PHRASES for exact matches: "multi-agent systems", "neural networks", "machine learning" - Combine related concepts with OR: "AI agents" OR "software agents" O...6 paramsSearch for papers on arXiv with advanced filtering and query optimization. QUERY CONSTRUCTION GUIDELINES: - Use QUOTED PHRASES for exact matches: "multi-agent systems", "neural networks", "machine learning" - Combine related concepts with OR: "AI agents" OR "software agents" O...
query*stringdate_tostringsort_bystringrelevance · datedate_fromstringcategoriesarraymax_resultsintegerdownload_paperDownload a paper from arXiv and return its full text content. Tries the HTML version first for clean extraction; falls back to PDF conversion if HTML is unavailable. Returns the paper content directly so you can read it immediately.1 paramsDownload a paper from arXiv and return its full text content. Tries the HTML version first for clean extraction; falls back to PDF conversion if HTML is unavailable. Returns the paper content directly so you can read it immediately.
paper_id*stringlist_papersList all papers that have been downloaded and stored locally via download_paper. Returns arXiv IDs only — use read_paper to access content. Returns an empty list if no papers have been downloaded yet. Workflow: search_papers -> download_paper -> list_papers -> read_paper.List all papers that have been downloaded and stored locally via download_paper. Returns arXiv IDs only — use read_paper to access content. Returns an empty list if no papers have been downloaded yet. Workflow: search_papers -> download_paper -> list_papers -> read_paper.
No parameters — call it with no arguments.
read_paperRead the full text content of a paper that was previously downloaded via download_paper. Returns the paper in markdown format. Will fail with a clear error if the paper has not been downloaded yet — call download_paper first. Workflow: search_papers -> download_paper -> read_p...1 paramsRead the full text content of a paper that was previously downloaded via download_paper. Returns the paper in markdown format. Will fail with a clear error if the paper has not been downloaded yet — call download_paper first. Workflow: search_papers -> download_paper -> read_p...
paper_id*stringget_abstractFetch the abstract and metadata of an arXiv paper by ID, WITHOUT downloading the full paper. Use this before download_paper to assess relevance and save tokens. Returns: title, authors, abstract, categories, published date, and PDF URL. Workflow tip: search_papers -> get_abstr...1 paramsFetch the abstract and metadata of an arXiv paper by ID, WITHOUT downloading the full paper. Use this before download_paper to assess relevance and save tokens. Returns: title, authors, abstract, categories, published date, and PDF URL. Workflow tip: search_papers -> get_abstr...
paper_id*stringsemantic_searchSemantic similarity search over papers you have already downloaded locally via download_paper. Supports free-text queries (e.g. 'attention mechanisms for long sequences') or finding papers similar to a given paper_id. IMPORTANT: only searches your local downloaded collection —...3 paramsSemantic similarity search over papers you have already downloaded locally via download_paper. Supports free-text queries (e.g. 'attention mechanisms for long sequences') or finding papers similar to a given paper_id. IMPORTANT: only searches your local downloaded collection —...
querystringpaper_idstringmax_resultsintegerreindexRebuild the local semantic index for downloaded papers.1 paramsRebuild the local semantic index for downloaded papers.
clear_existingbooleancitation_graphReturn papers citing an arXiv paper and papers that it references using Semantic Scholar's citation graph.1 paramsReturn papers citing an arXiv paper and papers that it references using Semantic Scholar's citation graph.
paper_id*stringwatch_topicSave or update a persistent research topic watch. When checked via check_alerts, returns only papers published since the last check — acting as a standing alert for new work on a topic. The topic string uses the same query syntax as search_papers (quoted phrases, field specifi...3 paramsSave or update a persistent research topic watch. When checked via check_alerts, returns only papers published since the last check — acting as a standing alert for new work on a topic. The topic string uses the same query syntax as search_papers (quoted phrases, field specifi...
topic*stringcategoriesarraymax_resultsintegercheck_alertsCheck all saved topic watches for newly published papers since the last check. Omitting the topic parameter runs ALL saved watches and returns new papers for each. Passing a topic string checks only that specific watch. Updates each watch's last_checked timestamp after running...1 paramsCheck all saved topic watches for newly published papers since the last check. Omitting the topic parameter runs ALL saved watches and returns new papers for each. Passing a topic string checks only that specific watch. Updates each watch's last_checked timestamp after running...
topicstringAn MCP server for searching arXiv, downloading papers, reading bounded full text, retrieving original LaTeX by section, following citation graphs, and maintaining research alerts.
It runs locally over stdio by default. Papers and indexes stay on your machine; search, source retrieval, citation graphs, and downloads call their respective external services.
The command-based integrations require uv, which provides uvx. Choose your client below; no repository clone or Python environment setup is required.
Add the MCP server for all projects:
claude mcp add --transport stdio --scope user arxiv \
-- uvx arxiv-mcp-server
For the richer plugin integration—which installs the MCP connection plus the bundled arXiv research skill—register this repository as a marketplace and install the plugin:
claude plugin marketplace add blazickjp/arxiv-mcp-server
claude plugin install arxiv-mcp-server@arxiv-mcp
Verify the direct MCP installation with claude mcp get arxiv. Restart Claude Code or run /reload-plugins after installing the plugin.
Add the MCP server:
codex mcp add arxiv -- uvx arxiv-mcp-server
Or install the MCP connection and bundled research skill as a Codex plugin:
codex plugin marketplace add blazickjp/arxiv-mcp-server
codex plugin add arxiv-mcp-server@arxiv-mcp
Verify the direct MCP installation with codex mcp get arxiv. Codex CLI, the Codex IDE extension, and Codex in the ChatGPT desktop app share this MCP configuration.
Use the Add to Kiro, Install in VS Code, or Install in VS Code Insiders button above.
For the richer Kiro Power integration, open the Powers panel, choose Add Custom Power → Import power from GitHub, and enter:
https://github.com/blazickjp/arxiv-mcp-server
The Power installs the MCP connection from mcp.json and adds focused arXiv research guidance. Kiro users who prefer manual configuration can place the generic configuration below in .kiro/settings/mcp.json for one workspace or ~/.kiro/settings/mcp.json for all workspaces.
macOS users can install a bundled .mcpb extension from the latest GitHub release:
arxiv-mcp-server-darwin-arm64-<version>.mcpbarxiv-mcp-server-darwin-x86_64-<version>.mcpbDouble-click the bundle, drag it into Claude Desktop, or open Settings → Extensions → Advanced settings → Install Extension…. The bundle includes the server dependencies and requires CPython 3.11.x.
Add this stdio configuration to any client that accepts standard MCP JSON:
{
"mcpServers": {
"arxiv": {
"type": "stdio",
"command": "uvx",
"args": ["arxiv-mcp-server"]
}
}
}
The default paper directory is ~/.arxiv-mcp-server/papers. To choose another directory, append "--storage-path", "/absolute/path/to/papers" to args.
For older papers that require PDF conversion, run the package with its PDF extra:
{
"mcpServers": {
"arxiv": {
"type": "stdio",
"command": "uvx",
"args": [
"--from",
"arxiv-mcp-server[pdf]",
"arxiv-mcp-server"
]
}
}
}
The supported package is published on PyPI. An unrelated npm package uses the same name, so do not install this server with npm, pnpm, or npx arxiv-mcp-server.
To place arxiv-mcp-server on your PATH instead of launching it through uvx:
uv tool install arxiv-mcp-server
Afterward, use "command": "arxiv-mcp-server" and omit the package name from args.
The repository now packages the same MCP server and research skill for both major plugin systems:
| Integration | Manifest | Marketplace |
|---|---|---|
| Claude Code | .claude-plugin/plugin.json | .claude-plugin/marketplace.json |
| OpenAI Codex / ChatGPT Work | .codex-plugin/plugin.json | .agents/plugins/marketplace.json |
| Kiro Power | POWER.md | mcp.json |
| Shared MCP launch | .mcp.json for Claude and repository-local clients; .codex-mcp.json for Codex plugins | uvx arxiv-mcp-server |
| Shared research workflow | skills/arxiv-mcp-server/SKILL.md | Installed with either plugin |
Direct MCP installation is the shortest path. Install the plugin when you also want the research workflow that steers the client toward focused searches, bounded reads, citation traversal, and section-level LaTeX retrieval.
The server currently exposes 14 tools.
| Tool | Purpose | Notes |
|---|---|---|
search_papers | Search arXiv by query, category, date, and sort order | Remote arXiv API |
get_abstract | Fetch metadata and an abstract by arXiv ID | Does not download the paper |
download_paper | Download and convert a paper to local Markdown | HTML first; PDF fallback uses [pdf] |
list_papers | List papers stored locally | Returns arXiv IDs |
read_paper | Read locally stored paper content | Supports start and max_chars |
get_paper_latex | Retrieve bounded author-submitted LaTeX | Remote arXiv source archive |
list_paper_latex_sections | Return a paginated LaTeX outline | Supports start and max_sections |
get_paper_latex_section | Read one bounded LaTeX section | Select by outline ID or exact title |
citation_graph | Fetch references and citing papers | Remote Semantic Scholar API |
export_citations | Export BibTeX for one or more arXiv IDs | Authoritative arXiv metadata |
watch_topic | Save or update an arXiv topic watch | Stored locally |
check_alerts | Check saved watches for new papers | Returns papers since the last check |
semantic_search | Search downloaded papers by semantic similarity | Requires [pro] |
reindex | Rebuild the local semantic index | Requires [pro] |
Ask your MCP client to call search_papers with:
{
"query": "\"Kolmogorov-Arnold Networks\"",
"categories": ["cs.LG", "cs.AI"],
"max_results": 5,
"sort_by": "date"
}
Then call get_abstract with:
{
"paper_id": "2404.19756"
}
Call download_paper with:
{
"paper_id": "2404.19756",
"max_chars": 12000
}
Then page through the cached content with read_paper:
{
"paper_id": "2404.19756",
"start": 0,
"max_chars": 12000
}
Large-content responses include content_length, returned_chars, next_start, and is_truncated. Pass next_start into the next call to continue reading.
Call get_paper_latex with:
{
"paper_id": "1706.03762"
}
Get the first page of its section outline with list_paper_latex_sections:
{
"paper_id": "1706.03762",
"start": 0,
"max_sections": 100
}
Then call get_paper_latex_section using an ID from that outline:
{
"paper_id": "1706.03762",
"section_id": "3.2",
"max_chars": 12000
}
LaTeX archives are validated, size-limited, and cached locally before content is returned.
Choose the install variant that matches the features you need:
# Base server
uv tool install arxiv-mcp-server
# Base server plus PDF conversion
uv tool install 'arxiv-mcp-server[pdf]'
# Base server plus local semantic search
uv tool install 'arxiv-mcp-server[pro]'
If the base tool is already installed, reinstall the selected variant:
uv tool install --force 'arxiv-mcp-server[pdf]'
The pdf extra installs pymupdf4llm and pymupdf-layout for papers without usable arXiv HTML. The pro extra adds local embedding dependencies for semantic_search and reindex; semantic search only operates on papers already downloaded to the configured storage directory.
The server provides seven MCP prompt workflows. Prompt availability depends on the client; the server provides workflow instructions but does not run a separate model.
| Prompt | Required arguments | Purpose |
|---|---|---|
research-discovery | topic | Map terminology, searches, papers, research clusters, and a reading path |
deep-paper-analysis | paper_id | Analyze one paper in depth |
summarize_paper | paper_id | Summarize methods, results, and limitations |
compare_papers | paper_ids | Compare multiple papers |
literature_review | topic | Synthesize a topic and optional paper set |
literature-synthesis | paper_ids | Synthesize themes, methods, timelines, or gaps across papers |
research-question | paper_ids, topic | Formulate grounded, falsifiable research questions |
For deployments where stdio is not practical:
TRANSPORT=http HOST=127.0.0.1 PORT=8080 \
uvx arxiv-mcp-server --storage-path /absolute/path/to/papers
Connect clients to:
{
"mcpServers": {
"arxiv": {
"type": "http",
"url": "http://127.0.0.1:8080/mcp"
}
}
}
The server binds to 127.0.0.1 by default and enables MCP DNS-rebinding protection. If a reverse proxy exposes the server, keep the process on a private interface and provide authentication and network controls upstream. Use ALLOWED_HOSTS and ALLOWED_ORIGINS for the host and origin values forwarded by the proxy.
| Setting | Default | Purpose |
|---|---|---|
--storage-path | ~/.arxiv-mcp-server/papers | Paper, source-cache, alert, and index storage |
MAX_RESULTS | 50 | Server-side cap for result counts |
REQUEST_TIMEOUT | 60 | PDF fallback download timeout in seconds |
TRANSPORT | stdio | stdio, http, or streamable-http |
HOST | 127.0.0.1 | HTTP bind host |
PORT | 8000 | HTTP bind port |
ALLOWED_HOSTS | empty | Additional accepted HTTP Host values |
ALLOWED_ORIGINS | empty | Additional accepted HTTP Origin values |
Environment variable names are case-insensitive through Pydantic settings. --storage-path is a command-line option rather than an environment setting.
Paper text and LaTeX are untrusted external content. A paper can contain text intended to manipulate an AI client into ignoring its instructions or calling unrelated tools.
See SECURITY.md for the reporting policy and threat details.
git clone https://github.com/blazickjp/arxiv-mcp-server.git
cd arxiv-mcp-server
uv sync --extra test --extra dev
uv run pytest
uv run black --check .
Run the development checkout from an MCP client with:
{
"mcpServers": {
"arxiv-dev": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/arxiv-mcp-server",
"run",
"arxiv-mcp-server"
]
}
}
}
Contributions are welcome. Read CONTRIBUTING.md before opening a pull request, and use GitHub Issues for reproducible bugs or scoped feature proposals.
Apache License 2.0. See LICENSE.
ARXIV_STORAGE_PATHOptional path for storing downloaded papers locally.
com.mcparmory/google-search
io.github.pipeworx-io/brave-search
marcopesani/mcp-server-serper
brave/brave-search-mcp-server
com.mcparmory/google-search-console
acamolese/google-search-console-mcp