Convert PDF to Markdown for RAG & LLMs
Feeding raw PDF text to an LLM throws away the structure your retrieval depends on. PDFjet returns clean Markdown — headings, lists and tables preserved in reading order — so you can chunk on real document structure and get better answers out of your RAG pipeline.
POST https://pdfjet.dev/extract/md with your PDF. You get chunk-ready Markdown back. Or call it straight from Claude/Cursor over MCP. Free API key, no card.
1. curl
curl -X POST https://pdfjet.dev/extract/md \
-H "Authorization: Bearer pj_live_..." \
-F [email protected] \
-o paper.md
2. Python (into your vector store)
import requests
md = requests.post(
"https://pdfjet.dev/extract/md",
headers={"Authorization": "Bearer pj_live_..."},
files={"file": open("paper.pdf", "rb")},
).text
# chunk on headings, then embed
chunks = [c for c in md.split("\n## ") if c.strip()]
3. Straight from your AI assistant (MCP)
PDFjet ships an MCP server, so Claude, Cursor and other MCP clients can convert a PDF to Markdown with no glue code:
{
"mcpServers": {
"pdfjet": {
"command": "npx",
"args": ["-y", "pdfjet-mcp"],
"env": { "PDFJET_API_KEY": "pj_live_..." }
}
}
}
Why Markdown beats a text dump
Flat pdftotext output merges columns, drops tables into noise, and loses heading boundaries — so your chunks straddle unrelated sections and retrieval degrades. Structured Markdown lets you split on #/## boundaries and keeps tables as tables, which the model can actually read.
See it on your own file first — preview the Markdown in the free tool, no signup.
Clean Markdown for your RAG pipeline
Free tier: 100 pages/month. No credit card.
Get your free API key →