Mistocr

Complete PDF OCR: clean markdown with structure and visual content for RAG and LLM pipelines

mistocr converts PDFs into clean markdown for RAG and LLM pipelines. It runs Mistral’s OCR, then uses an LLM to fix the heading hierarchy and to describe the images in the document.

OCR is often treated as a trivial step in AI pipelines. In practice, a poorly converted PDF gives every downstream step poor input. mistocr came out of months of processing large collections of niche-domain PDFs. It handles three problems that raw OCR output leaves:

mistocr uses Mistral’s mistral-ocr-latest model for OCR. It uses fastllm for the LLM steps, which supports Anthropic, OpenAI and Gemini models.

Note

This tutorial shows mistocr on a real PDF. It also shows how the fixed headings let you navigate a long document by its structure.

Get Started

Install mistocr from PyPI. It needs Python 3.10 or later.

$ pip install mistocr

Set your API keys as environment variables. The OCR step needs MISTRAL_API_KEY. The heading and image steps need the key for your LLM provider: ANTHROPIC_API_KEY for the default Claude model, OPENAI_API_KEY for OpenAI models, or GEMINI_API_KEY for Gemini models.

import os
os.environ['MISTRAL_API_KEY'] = 'your-key-here'
os.environ['ANTHROPIC_API_KEY'] = 'your-key-here'

Complete Pipeline

Single file processing

Run OCR, heading fixes and image descriptions on one PDF:

from mistocr.pipeline import pdf_to_md
await pdf_to_md('files/test/resnet.pdf', 'files/test/md_test')
mistocr.pipeline - INFO - Step 1/3: Running OCR on files/test/resnet.pdf...
mistocr.core - INFO - Waiting for batch job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 (initial status: QUEUED)
mistocr.core - DEBUG - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 status: QUEUED (elapsed: 0s)
mistocr.core - DEBUG - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 status: RUNNING (elapsed: 2s)
mistocr.core - DEBUG - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 status: RUNNING (elapsed: 2s)
mistocr.core - DEBUG - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 status: RUNNING (elapsed: 2s)
mistocr.core - INFO - Job 4ec899ca-ada8-4fa7-8894-0191ff6ac4e5 completed with status: SUCCESS
mistocr.pipeline - INFO - Step 2/3: Fixing heading hierarchy...
mistocr.pipeline - INFO - Step 3/3: Adding image descriptions...
Describing 12 images...
mistocr.pipeline - INFO - Done!
Saved descriptions to /tmp/tmp62c7_ac1/resnet/img_descriptions.json
Adding descriptions to 12 pages...
Done! Enriched pages saved to files/test/md_test

pdf_to_md does the following:

  1. Runs OCR on the PDF with Mistral’s OCR API.
  2. Fixes the heading hierarchy.
  3. Describes the images, such as charts and diagrams, and adds the descriptions to the markdown.
  4. Saves the result to files/test/md_test.

The output folder looks like this:

files/test/md_test/
├── img/
│   ├── img-0.jpeg
│   ├── img-1.jpeg
│   └── ...
├── page_1.md
├── page_2.md
└── ...

Each image link in a page is followed by its description:

![Figure 1](img/img-0.jpeg)
AI-generated image description:
___
A residual learning block...
___

read_pgs reads the processed document back as one markdown string:

from mistocr.core import read_pgs
md = read_pgs('files/test/md_test')
print(md[:500])
# Deep Residual Learning for Image Recognition  ... page 1

Kaiming He Xiangyu Zhang Shaoqing Ren Jian Sun<br>Microsoft Research<br>\{kahe, v-xiangz, v-shren, jiansun\}@microsoft.com


## Abstract ... page 1

Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, ins

read_pgs joins all pages by default. Pass join=False to get a list of pages. Each heading ends with ... page N, the page it comes from.

Advanced Usage

Batch OCR for entire folders:

from mistocr.core import ocr_pdf

# OCR all PDFs in a folder using Mistral's batch API
output_dirs = ocr_pdf('path/to/pdf_folder', dst='output_folder')

Custom models and prompts for heading fixes:

from mistocr.refine import fix_hdgs

# Use a different model or custom prompt
await fix_hdgs('ocr_output/doc1', 
               model='gpt-4o',
               prompt=your_custom_prompt)

Custom image description with rate limiting:

from mistocr.refine import add_img_descs

# Use a different model, and change how fast images are sent
await add_img_descs('ocr_output/doc1',
                    model='claude-opus-4-6',
                    semaphore=5,  # concurrent requests, default 2
                    delay=0.5)    # seconds to wait after each request, default 1

For complete control over each pipeline step, see the core, refine, and pipeline module documentation.

Developer Guide

mistocr is developed with nbdev. If you’re new to nbdev, start with its getting started guide.

Install mistocr in Development mode

# install mistocr in development mode
$ pip install -e .

# edit the notebooks in nbs/

# export the notebooks to mistocr, run the tests and clean the notebooks
$ nbdev-prepare

Documentation

The documentation is on the repository’s GitHub Pages. The package is on PyPI.