You can turn a PDF into LLM-ready Markdown by uploading it to a service, by pasting its text into a chat, or by asking your agent to call a conversion engine on your own machine. The third route keeps headings, tables and lists as real Markdown and uploads nothing. With the GroupDocs.Conversion.Mcp server, Claude, Cursor or GitHub Copilot run it locally from one prompt:
Convert report.pdf to Markdown
The step-by-step version with config and troubleshooting is in the documentation: How to convert PDF to Markdown with an MCP server.
Why does the Markdown step decide the quality of your RAG index?
A retrieval pipeline can only index what the conversion step kept. If a table is flattened into a run of numbers, a chunk of that table answers no question. If headings disappear, a chunker has no section boundaries to respect. The documents you most want in a private knowledge base are often the ones you least want to send to a third party: contracts, internal reports, patient or customer records. The choice of route therefore decides two things at once: how much structure survives, and where the file travels.
Way 1: Do you upload the PDF to an online converter or a cloud parsing API?
This is the quickest route and the one with the largest trade-off. The file leaves your machine and is processed on infrastructure you do not control. For public documents that may be acceptable. For anything confidential it is the step a security review will stop.
Choose it only when the content is not sensitive and you accept a per-file round trip through someone else’s service.
Way 2: Do you paste the PDF text into the chat and ask the model to format it?
This keeps the process on your side of the table, but the model is now rewriting the content, not converting it. It sees only the text that was extracted when you pasted. Table borders, column order and page geometry are already gone, and the model fills the gaps by guessing. Long documents also hit the context window before you reach the last page.
Choose it for a one-off paragraph, not for a folder of reports.
Way 3: Do you let your agent call a local conversion engine?
Here the agent decides and the engine converts. The agent sends a file name and a target format to the convert tool, and the GroupDocs.Conversion engine writes a Markdown file next to your documents. In the tool call, the format is md:
{
"name": "convert",
"arguments": {
"file": { "filePath": "report.pdf" },
"format": "md"
}
}
You never write that JSON. You type the sentence, and the agent produces the call. Structure is preserved as real Markdown constructs: headings become # and ##, tables become pipe tables, lists become list items. That is the format an embedding pipeline and an LLM read natively, and you can feed the file straight to your chunker.
The shape of the result looks like this. The content below is illustrative, not output from a real run:
# Quarterly report
## Revenue by region
| Region | Q1 | Q2 |
|---|---|---|
| North | 120 | 135 |
| South | 98 | 101 |
Complex layouts are where an engine built for document fidelity differs from a basic text extractor. Whether it matters for your documents is something to test on your three hardest PDFs before you commit.
If Markdown is the only output you need, GroupDocs also has a dedicated server, GroupDocs.Markdown.Mcp. Its documentation describes it as turning PDF, Word, Excel, EPUB and 20+ more formats into clean, structured Markdown, with control over page selection, images and front matter. GroupDocs.Conversion.Mcp fits when Markdown is one of many target formats in the same workflow. The RAG-focused comparison is in Your RAG pipeline starts with Markdown — keep that step local, and the server is documented at docs.groupdocs.com/markdown/mcp.
How do you set up the local route?
Register the server in your client once. A Claude Desktop entry:
{
"mcpServers": {
"groupdocs-conversion": {
"type": "stdio",
"command": "dnx",
"args": ["GroupDocs.Conversion.Mcp", "--yes"],
"env": { "GROUPDOCS_MCP_STORAGE_PATH": "/path/to/documents" }
}
}
}
This needs the .NET 10 SDK. The same server entry works in Cursor, VS Code with GitHub Copilot and Claude Code; the exact config file for each client is listed in the docs hub. Restart the client afterwards. In evaluation or license-file mode, conversion needs no network once the package is downloaded; metered licensing reports usage and so needs outbound access.
Prompts that work:
Turn every PDF in this folder into Markdown for my knowledge base.
Check how many pages whitepaper.pdf has, then convert it to Markdown.
The second prompt chains get_document_info, which returns file type, page count and basic properties, before convert.
Example session
This is an abridged, illustrative session based on the tool descriptions in the documentation, not a screenshot.
You: Check how many pages whitepaper.pdf has, then convert it to Markdown.
Agent: [calls get_document_info with file.filePath = "whitepaper.pdf"]
Agent: [calls convert with file.filePath = "whitepaper.pdf", format = "md"]
Agent: Done. The Markdown file is saved as whitepaper.md in your output folder.
What are the limits?
- Scans need OCR first. An image-only PDF has no text layer, and OCR is not part of this server, so the Markdown will not contain its words. Run OCR before conversion if your PDFs are scans.
- Evaluation mode. Without a license, output carries an evaluation watermark and one server process can open at most 15 documents. Call
get_license_statusbefore a large ingestion to confirm which mode is active. - Conversion, not extraction. This server converts a whole document to Markdown. If you need specific fields or typed values pulled out, use GroupDocs.Parser.Mcp (delivered as a Docker image only), and use both when a pipeline needs both.
FAQ
How do I convert PDF to Markdown for RAG without uploading the file?
Run a conversion server on your own machine and let your agent call it. With GroupDocs.Conversion.Mcp the convert tool writes the .md file to your folder, and no document is sent to a cloud service.
Are tables kept when a PDF becomes Markdown? Yes, as Markdown pipe tables. Test your most complex tables first, because layout complexity is where converters differ.
Can I convert the Markdown back to PDF or Word later? Yes. The same server converts Markdown to PDF and DOCX, which is covered in the export post.
Go deeper
- Documentation, canonical how-to: How to convert PDF to Markdown with an MCP server
- Documentation hub: GroupDocs.Conversion MCP Server
- Start here: Why your AI agent should not convert documents by itself
- Related: Automate report delivery: let the agent write Markdown and ship Word
- Related: A whole folder, one prompt: batch conversion with Claude and MCP
- Announcement: Your AI Agent Can Now Convert Documents — Locally. Introducing GroupDocs.Conversion MCP
- On-premise and security model: 3 architectures for AI document processing, and the one that keeps files inside your network
- Questions: GroupDocs Conversion forum
- Source: GroupDocs.Conversion.Mcp on GitHub