Complete PDF OCR: clean markdown with structure and visual content for RAG and LLM pipelines
mistocr converts PDFs into clean markdown for RAG and LLM pipelines. It runs Mistral’s OCR, then uses an LLM to fix the heading hierarchy and to describe the images in the document.
OCR is often treated as a trivial step in AI pipelines. In practice, a poorly converted PDF gives every downstream step poor input. mistocr came out of months of processing large collections of niche-domain PDFs. It handles three problems that raw OCR output leaves:
Heading hierarchy. OCR often gets heading levels wrong in long documents. mistocr sends all the headings to an LLM, which returns the corrected levels.
Visual content. mistocr classifies each image as informative or decorative, and adds a description of it to the markdown. Charts, figures and diagrams become searchable text.
Cost. mistocr sends a single PDF straight to Mistral’s OCR API. It sends several PDFs through Mistral’s batch API, which costs half as much: $1 instead of $2 per 1,000 pages.
mistocr uses Mistral’s mistral-ocr-latest model for OCR. It uses fastllm for the LLM steps, which supports Anthropic, OpenAI and Gemini models.
Note
This tutorial shows mistocr on a real PDF. It also shows how the fixed headings let you navigate a long document by its structure.
Get Started
Install mistocr from PyPI. It needs Python 3.10 or later.
$ pip install mistocr
Set your API keys as environment variables. The OCR step needs MISTRAL_API_KEY. The heading and image steps need the key for your LLM provider: ANTHROPIC_API_KEY for the default Claude model, OPENAI_API_KEY for OpenAI models, or GEMINI_API_KEY for Gemini models.
Saved descriptions to /tmp/tmp62c7_ac1/resnet/img_descriptions.json
Adding descriptions to 12 pages...
Done! Enriched pages saved to files/test/md_test
read_pgs reads the processed document back as one markdown string:
from mistocr.core import read_pgsmd = read_pgs('files/test/md_test')print(md[:500])
# Deep Residual Learning for Image Recognition ... page 1
Kaiming He Xiangyu Zhang Shaoqing Ren Jian Sun<br>Microsoft Research<br>\{kahe, v-xiangz, v-shren, jiansun\}@microsoft.com
## Abstract ... page 1
Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, ins
read_pgs joins all pages by default. Pass join=False to get a list of pages. Each heading ends with ... page N, the page it comes from.
Advanced Usage
Batch OCR for entire folders:
from mistocr.core import ocr_pdf# OCR all PDFs in a folder using Mistral's batch APIoutput_dirs = ocr_pdf('path/to/pdf_folder', dst='output_folder')
Custom models and prompts for heading fixes:
from mistocr.refine import fix_hdgs# Use a different model or custom promptawait fix_hdgs('ocr_output/doc1', model='gpt-4o', prompt=your_custom_prompt)
Custom image description with rate limiting:
from mistocr.refine import add_img_descs# Use a different model, and change how fast images are sentawait add_img_descs('ocr_output/doc1', model='claude-opus-4-6', semaphore=5, # concurrent requests, default 2 delay=0.5) # seconds to wait after each request, default 1
For complete control over each pipeline step, see the core, refine, and pipeline module documentation.
# install mistocr in development mode$ pip install -e .# edit the notebooks in nbs/# export the notebooks to mistocr, run the tests and clean the notebooks$ nbdev-prepare