by jztan
An MCP server that gives your AI agent agentic RAG over your PDFs, one file or a whole folder: hybrid semantic + keyword search, selective page reads, tables, images, OCR, chart data, and multi-column/CJK layouts. The agent decides when to search; pdf-mcp does the retrieval.
# Add to your Claude Code skills
git clone https://github.com/jztan/pdf-mcpLast scanned: 8/9/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-08-09T05:04:04.247Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}See how pdf-mcp compares with popular alternatives.
pdf-mcp is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by jztan. An MCP server that gives your AI agent agentic RAG over your PDFs, one file or a whole folder: hybrid semantic + keyword search, selective page reads, tables, images, OCR, chart data, and multi-column/CJK layouts. The agent decides when to search; pdf-mcp does the retrieval. It has 134 GitHub stars.
Yes. pdf-mcp passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/jztan/pdf-mcp" and add it to your Claude Code skills directory (see the Installation section above).
pdf-mcp is primarily written in Python. It is open-source under jztan on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh pdf-mcp against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
Agentic RAG over your PDFs, one file or a whole folder, as a single MCP tool.
The agent decides when to search; pdf-mcp does the retrieval and hands back excerpts. It is an MCP server that lets Claude Code and other AI agents search one PDF or a whole folder by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts, with optional CUDA acceleration for warming large corpora.
mcp-name: io.github.jztan/pdf-mcp
Drop in any PDF, or a whole folder of them, and watch an agent triage the corpus, search across every document at once, and read only the pages that matter, using a fraction of the tokens. 100% client-side, no install required.
| Without pdf-mcp | With pdf-mcp | |
|---|---|---|
| Large PDFs | Context overflow | Read only the pages you need |
| Finding content | Load everything | Hybrid search: BM25 keyword + semantic |
| Folders of PDFs | One document at a time | Warm, triage, and search a whole folder |
| Warming a big folder | Minutes of CPU embedding | Length-sorted small-batch CPU encode; optional CUDA embedding, one to two orders of magnitude faster on an NVIDIA card |
| Tables and charts | Lost in raw text | Structured rows, and (x, y) data from vector charts |
| Multi-column and vertical layouts | Columns interleaved | Correct reading order, including Japanese tategaki |
| Scanned PDFs | No text at all | OCR via Tesseract, parallel across pages |
| Repeated access | Re-parse every time | SQLite cache that survives restarts |
| Hidden or injected text | Silently ingested | Flagged as untrusted, nothing stripped |
Coming in the next release. The download link works once that release is out; until then, use the terminal install below.
The first start downloads pdf-mcp's components (about 250 MB) and can take a few minutes; later starts take seconds. OCR for scanned pages is included: the first scanned page downloads an English-only Tesseract (about 14 MB). Needs Windows 10 or later, or macOS 13 or later (14 on Apple Silicon), and works in Claude Desktop's Chat. Updating, uninstalling and other details are in docs/clients.md.
Needs Python 3.10 or later. Install the pdf-mcp command with
uv or pipx:
uv tool install pdf-mcp # or: pipx install pdf-mcp
pip install pdf-mcp works inside a virtual environment; Homebrew's Python
and recent Debian and Ubuntu refuse a system-wide pip install.
Then add it to your client. For Claude Code:
claude mcp add pdf-mcp -- pdf-mcp
For VS Code, Cursor, Codex CLI, Kiro or any other MCP client, see docs/clients.md. Then ask your agent to read a PDF.
Search, the corpus tools, tables and multi-column and CJK reading order work out of the box. OCR on scanned pages also needs Tesseract:
brew install tesseract # macOS
sudo apt install tesseract-ocr # Ubuntu/Debian
winget install -e --id UB-Mannheim.TesseractOCR # Windows
Optional: CUDA embedding on an NVIDIA card warms large folders one to two orders of magnitude faster; see docs/configuration.md.
pdf-mcp's tools are also plain Python functions, so you can import them and hand a PDF to the Anthropic SDK without running a server. Two runnable scripts, for a question and for a whole document: examples/.
Why this exists, and what broke along the way: Claude's 100-page PDF limit and how I got around it
13 specialized tools rather than one monolithic one. Typical pattern:
pdf_info to plan, pdf_search to locate (its paragraph excerpts often
answer the question outright), pdf_read_pages when you need more. For a
folder, pdf_corpus_overview to triage, then pdf_corpus_search.
| Tool | What it does |
|---|---|
pdf_info |
Page count, metadata, TOC summary, scanned-page detection. Call first. |
pdf_search |
Hybrid search (keyword + semantic), page or section granularity, paragraph or context-window excerpts with source coordinates |
pdf_read_pages |
Read specific pages or ranges, with OCR on demand, tables, and embedded images |
pdf_read_all |
Read a whole document in one call, byte-capped |
pdf_get_toc |
Full table of contents for documents with many bookmarks |
pdf_render_pages |
Render pages as PNG for vision models: diagrams, handwriting, scans |
pdf_extract_chart |
Chart data as exact (x, y) tables, read from plot geometry |
pdf_corpus_warm |
Warm a folder of PDFs into the cache within a time budget |
pdf_corpus_overview |
Per-document triage cards for a folder |
pdf_corpus_search |
Search across a folder, with document and page provenance; excerpt_style="auto" picks the excerpt unit per query |
pdf_cache_stats |
Per-document cache breakdown and total size |
pdf_cache_clear |
Clear expired or all cache entries |
server_info |
Which optional features and config are active |
Text returned by any of these is untrusted content extracted from a PDF.
pdf_info(content_trust=True) reports hidden text a human reader cannot
see, and the read tools flag it per page.
Example prompts:
"Read the PDF at /path/to/document.pdf"
"Which pages discuss supply chain risks?"
"Find sections about the training process"
"Show me what page 5 looks like"
"OCR pages 3-5 of the scanned PDF"
Full reference, every parameter and response shape: docs/tool-reference.md. Embedding model selection: docs/embedding-models.md.
For a large document (e.g., a 200-page annual report):
User: "Summarize the risk factors in this annual report"
Agent workflow:
1. pdf_info("report.pdf")
→ 200 pages, TOC shows "Risk Factors" on page 89
2. pdf_search("report.pdf", "risk factors")
→ Matches with structural paragraph excerpts: each excerpt
is the bullet, paragraph, or heading that matched, not a
fixed-width window. Often enough to answer directly.
3. If excerpts are sufficient → synthesize answer
4. If more context needed:
pdf_read_pages("report.pdf", "89-95")
→ Full page text for deeper reading
STDIO is the default and is what every example above uses. pdf-mcp-http
serves the same tools over HTTP, for clients that cannot spawn a process
(the Anthropic API MCP connector, claude.ai custom connectors) and for a
warm corpus shared by several clients.
export PDF_MCP_AUTH_TOKEN="$(openssl rand -hex 32)"
pdf-mcp-http
Paths resolve on the server, so an HTTP agent reads what is already there:
files under an allow-listed root, or a URL the server fetches. It cannot
hand over a file from its own machine. It is single-tenant and fails
closed: with no auth token and no [paths] allow list, the process exits
rather than serving an open endpoint.
Docker images are published to GHCR for amd64 and arm64, with everything baked in, so every tool works on the first request:
./deploy.sh # token, image, start, health-check
cp your.pdf documents/ # this folder is the server's /data/pdfs
Read docs/remote-access.md for the trust boundary and threat model before deploying, and docs/configuration.md for setup, client config, and token rotation.
pdf-mcp works out of the box. To restrict which paths and URL hosts the server may touch, tune cache and worker settings, or add your own content-trust phrases, see docs/configuration.md.
See ROADMAP.md for planned features and release history.
Contributions are welcome. See docs/contributing.md for setup, checks,