Skip to content

AI Environment

PDF to Markdown Conversion Guide

Convert PDF research papers into clean Markdown so AI tools can read, quote, and analyze them more reliably.

Intermediate · 20–30 min · MinerU or Mistral API key, depending on method

PDFs were built for human eyes, not machine reasoning.

Overview

You need clean text, not raw PDFs, for effective AI research workflows. Raw PDFs often contain complex layouts, images, and formatting that can interfere with AI processing, embeddings, and analysis. This guide covers multiple methods for converting PDFs to clean, AI-friendly text formats.

Benefits:

  • Better AI processing: Clean text without layout artifacts or formatting issues
  • Cost efficiency: Pre-convert PDFs once instead of processing them repeatedly
  • Token optimization: Text formats use fewer tokens than PDF processing
  • Embedding quality: Consistent text extraction improves embedding accuracy
  • Workflow integration: Text files work with all AI tools and models

Common use cases:

  • Research literature analysis
  • AI-assisted content summarization
  • Embedding creation for semantic search
  • Large language model context preparation

For a born-digital report with one column, pdftotext from Poppler already does the job from the command line, and any agent can call it. Academic papers are the hard case: two columns, footnotes, running heads, figures with captions, tables, and appendices. A plain text dump interleaves the columns and loses which figure a sentence was about. The methods below are layout-aware for that reason, and the ones that keep the figures and a layout file let a model read a paper's structure without relying on its own vision, which not every model has in equal measure.

PDF conversion methods

Method 1: MinerU MCP and CLI (cloud-backed)

Use for agent workflows and scripted corpus preparation when cloud processing is appropriate.

MinerU MCP exposes parsing to an agent; mineru-cloud exposes the same tools to a shell. Both use MinerU's hosted API, including when the input is a local file. See the MinerU guide for package availability, tool schemas, model options, and setup.

Quick setup:

claude mcp add mineru-mcp -e MINERU_API_KEY=your-key -- npx -y mineru-mcp

Then ask Claude: "Parse this PDF with VLM mode: [URL]"

For a long document, use parse-long followed by merge-slices: CLI example. Supply the page count and inspect the merged slice count, notes, figures, and page references. Page-range selection is not local redaction: the local-file route uploads the entire document once per slice. The limits note records the discrepancy between the source owner's observed cap and the vendor guide.

Pros: agent and shell interfaces, batch submission, structured output with available figures. Cons: requires an API key and cloud-upload permission; extraction still needs checks against the original. CLI help/list do not parse or upload documents.


Method 2: MinerU desktop client and web app (free, cloud-backed)

Use for one-off conversions and for participants who have not set up an agent yet.

MinerU's desktop client and web app run the same parser as the API, on MinerU's own compute. Drag a PDF in, or paste the URL of a publicly available one, and the result comes back as a folder: the Markdown, an images directory the Markdown links to, and layout files (content_list, model, middle JSON) that map text, tables, and figures to their positions on the page. The layout files are what let a model without strong vision still read the paper's structure.

  1. Visit MinerU and download the client for Windows, macOS, or Linux, or use the web version.
  2. Drop the PDF in, or enter its URL, and start the conversion.
  3. Save the output folder next to the paper. Convert once; every later agent reads the Markdown.

Nothing installs beyond the client and no API key is needed, but the file does leave your machine, and usage is quota-gated rather than metered; check MinerU's docs for the current daily limits. For a fully local run with no upload, install the open-source package instead (uv pip install -U "mineru[all]", source at GitHub). Its pipeline backend runs on CPU (mineru -p paper.pdf -o out -b pipeline), slower and less accurate than the VLM backend, which wants a capable GPU. Confidential papers take the local route; the cloud client is a convenience, not the only option. Pros: free, no key, no agent required, keeps figures and layout. Cons: uploads the file; large batch jobs are better served by Method 1 or 3.


Method 3: Mistral OCR script (hosted batch)

Use for scripted batch jobs through Mistral's hosted OCR API.

Use the Mistral OCR script approach for bulk processing outside of Claude (100+ papers at once).

This is another cloud route, not offline OCR. Choose between it and MinerU based on the documents, approved data handling, and results from a small comparison; a locally running script still sends document data to its provider.


Method 4: manual copy-paste (fallback)

Use for emergency single documents when other methods are unavailable.

  1. Open the PDF in a PDF reader.
  2. Select and copy the text from each page.
  3. Paste the text into a text editor or Markdown file.
  4. Save the file with a .md or .txt extension.

Limitations: time-consuming, layout issues, manual errors. Use MinerU instead.

Mistral OCR API setup (optional)

Use Mistral OCR when you want hosted OCR that preserves document layout and outputs Markdown. The maintained script is scripts/ocr/mistral_batch_ocr.py.

Step 1: get Mistral API key

  1. Visit the Mistral AI Console.
  2. Create an account or sign in.
  3. Navigate to the API Keys section.
  4. Click "Create new key".
  5. Copy and save your API key securely.

Check current pricing and limits in the Mistral console before large batches.

Step 2: environment setup

Set your API key as an environment variable:

export MISTRAL_API_KEY="your_api_key_here"

Step 3: run the batch OCR script

pip install mistralai
python scripts/ocr/mistral_batch_ocr.py readings/pdfs readings/markdown

The script:

  • Recursively processes PDFs and preserves folder structure
  • Uses Mistral batch jobs against /v1/ocr
  • Writes page.markdown from current OCR responses and falls back to older page.text shapes
  • Cleans up request/result JSONL files unless you set --keep-work-files
  • Skips large PDFs above --max-size-mb
# Include extracted image payloads if needed
python scripts/ocr/mistral_batch_ocr.py readings/pdfs readings/markdown --include-images

# Keep JSONL files for debugging
python scripts/ocr/mistral_batch_ocr.py readings/pdfs readings/markdown --keep-work-files

Local Baidu/PaddleOCR option

Use local PaddleOCR when you need high-volume OCR without metered hosted API calls. PaddleOCR is Baidu's open-source OCR toolkit. Current model families such as PP-OCRv5 and PP-OCRv6 can run locally after you download the model and runtime.

The maintained script is scripts/ocr/paddle_unlimited_ocr.py.

pip install paddleocr
# Install PaddlePaddle for your platform from the official Paddle install page.

python scripts/ocr/paddle_unlimited_ocr.py readings/pdfs readings/markdown --lang en
python scripts/ocr/paddle_unlimited_ocr.py readings/pdfs readings/markdown --lang ch --ocr-version PP-OCRv5

Use this route for:

  • Sensitive documents that should not leave the machine
  • Very large batches where hosted OCR cost is the constraint
  • Chinese or multilingual OCR experiments where PaddleOCR models are a good fit

Avoid the local PaddleOCR route when you need turnkey setup. Local runtime and model installation can be less reliable than a hosted API.

Alternative API options

Google Document AI

Use Google Document AI when you already work in Google Cloud.

  1. Visit Google Cloud Console
  2. Enable the Document AI API.
  3. Create a processor for OCR.
  4. Use the Python client library for batch processing.

Azure Form Recognizer

Use Azure Document Intelligence in enterprise Azure environments.

  1. Visit Azure Portal
  2. Create a Cognitive Services resource.
  3. Use the Form Recognizer service.
  4. Integrate using the REST API or SDK.

Best practices and tips

File organization

your-project/
├── pdfs/           # Original PDFs
│   ├── session-1/
│   ├── session-2/
│   └── articles/
└── markdown/       # Converted text files
    ├── session-1/
    ├── session-2/
    └── articles/

Quality control

  • Spot check: review converted files for accuracy
  • Complex layouts: Some academic PDFs may need manual review
  • Images and tables: OCR may not capture complex visual elements
  • Languages: Ensure the API supports your document languages

Cost optimization

  • Batch processing: Convert all PDFs at once rather than individually
  • File size limits: Be aware of API size restrictions
  • Free tiers: Use free options for small projects
  • Caching: Store converted files to avoid re-processing

Workflow integration

  • Version control: Track both original PDFs and converted text
  • Backup: Keep original PDFs as source of truth
  • Naming: Maintain consistent file naming conventions
  • Metadata: Preserve citation information alongside converted text

Troubleshooting

Common issues

"API Key Not Found"

  • Verify environment variable is set: echo $MISTRAL_API_KEY
  • Restart your terminal or command prompt
  • Check for typos in variable name

"File Too Large"

  • The default limit is 36 MB per PDF
  • Split large documents or use the --max-size parameter
  • Consider alternative conversion methods for very large files

"Processing Failed"

  • Check the PDF file integrity
  • Some PDFs may have copy protection
  • Try alternative conversion methods

"Rate Limits Exceeded"

  • APIs have rate limits (requests per minute/hour)
  • Implement delays between requests
  • Consider paid plans for higher limits

Integration with case study workflow

  • Initial setup: convert all PDFs at project start
  • Ongoing: convert new PDFs as they're added
  • Processing: use converted Markdown files for all AI workflows
  • Storage: maintain both original PDFs and converted text

Case study references

  • AI-assisted literature analysis
  • Human-AI synthesis workflows
  • Agentic workflow design
  • API Keys Guide: For Mistral setup
  • Model Reference Guide: For compatible AI models

Advanced options

Next steps

  1. Choose your preferred conversion method.
  2. Convert a few sample PDFs.
  3. Process your full document collection.
  4. Use converted files in your AI workflows.
  5. Track conversion quality and costs.

Converting PDFs once avoids the time and cost of processing them repeatedly in later AI workflows.

cite this page

Lin, X. (2026). PDF to Markdown Conversion Guide. Research Memex. https://research-memex.org/docs/implementation/ai-environment-setup/ocr-pdf-conversion-guide

@misc{docs-implementation-ai-environment-setup-ocr-pdf-conversion-guide-2026,
  author = {Xule Lin},
  title = {PDF to Markdown Conversion Guide},
  year = {2026},
  howpublished = {\url{https://research-memex.org/docs/implementation/ai-environment-setup/ocr-pdf-conversion-guide}},
  note = {ORCID: 0000-0001-7885-4194}
}

one renderingthe source remains