get_api_key
def get_api_key(
key:str=None, # Mistral API key
):Get Mistral API key from parameter or environment
This module provides OCR processing for PDF documents using Mistral’s OCR API. For single PDFs, it uses direct (synchronous) OCR for immediate results. For multiple PDFs, it uses the batch API for cost savings. Results are saved as markdown files with extracted images. The main entry point is ocr_pdf which auto-detects the best mode based on input.
To use this module, you’ll need: - A Mistral API key set in your environment as MISTRAL_API_KEY or passed to functions - PDF files to process - The mistralai Python package installed
Load your API key from a .env file or set it directly:
Get Mistral API key from parameter or environment
For processing a single PDF with immediate results (no batch queue), we use direct OCR via base64 encoding.
Run direct OCR on a single PDF (no batch queue)
(15,
'# Attention Is All You Need\n\nAshish Vaswani*\n\nGoogle Brain\n\navaswani@google.com\n\nNoam Shazeer*\n\nGoogle Brain\n\nnoam@google.com\n\nNiki Parmar*\n\nGoogle Research\n\nnikip@google.com\n\nJakob Uszkoreit*\n\nGoogle')
The pipeline works in stages: first we upload PDFs to Mistral’s servers and get signed URLs, then we create batch entries for each PDF, submit them as a job, monitor completion, and finally download and save the results as markdown files with extracted images.
Upload PDF to Mistral and return signed URL
('https://mistralaifilesapiprodswe.blob.core.windows.net/fine-tune/d3c29c82-04d2-4c77-98b6-d2820e375e1f/75695561-4561-4dc1-aeab-b0f34f40f753/b115c933188b4371a6a11793c2a59ace.pdf?se=2026-02-07T12%3A23%3A37Z&sp=r&sv=2026-02-06&sr=b&sig=rtxkyBXX6O3eyLE0Z8Zw3LuS6AM%2BWqpoAnegHDeSVsM%3D',
<mistralai.sdk.Mistral>)
Each PDF needs a batch entry that tells Mistral where to find it and what options to use. The custom id cid helps you identify results later - by default it uses the PDF filename.
def create_batch_entry(
path:str, # Path to PDF file,
url:str, # Mistral signed URL
cid:str=None, # Custom ID (by default using the file name without extension)
inc_img:bool=True, # Include image in response
extract_header:bool=True, # Extract headers from document
extract_footer:bool=True, # Extract footers from document
)->dict[str, str | dict[str, str | bool]]: # Batch entry dictCreate a batch entry dict for OCR
When extract_header and extract_footer are True, headers and footers are extracted as separate page properties and excluded from the markdown output. Mistral defaults both to False, which includes them in the markdown.
{'custom_id': 'attention-is-all-you-need',
'body': {'document': {'type': 'document_url',
'document_url': 'https://mistralaifilesapiprodswe.blob.core.windows.net/fine-tune/d3c29c82-04d2-4c77-98b6-d2820e375e1f/75695561-4561-4dc1-aeab-b0f34f40f753/b115c933188b4371a6a11793c2a59ace.pdf?se=2026-02-07T12%3A23%3A37Z&sp=r&sv=2026-02-06&sr=b&sig=rtxkyBXX6O3eyLE0Z8Zw3LuS6AM%2BWqpoAnegHDeSVsM%3D'},
'include_image_base64': True,
'extract_header': True,
'extract_footer': True}}
Upload PDF and create batch entry in one step
{'custom_id': 'attention-is-all-you-need',
'body': {'document': {'type': 'document_url',
'document_url': 'https://mistralaifilesapiprodswe.blob.core.windows.net/fine-tune/d3c29c82-04d2-4c77-98b6-d2820e375e1f/75695561-4561-4dc1-aeab-b0f34f40f753/b115c933188b4371a6a11793c2a59ace.pdf?se=2026-02-07T15%3A38%3A48Z&sp=r&sv=2026-02-06&sr=b&sig=U%2Bzo%2BntyDgW4BKNG2T7hP/OYDsSaUFtZMyydxugOXio%3D'},
'include_image_base64': True,
'extract_header': True,
'extract_footer': True}}
Once you have batch entries for all your PDFs, submit them as a single job. Mistral processes them asynchronously, so you’ll need to poll for completion.
Submit batch entries and return job
BatchJobOut(id='482ba9ee-559f-4b11-8f60-8aabe739ad05', input_files=['ce74af1e-a2b4-4d8e-9256-fab329c3d2de'], endpoint='/v1/ocr', errors=[], status='QUEUED', created_at=1770393138, total_requests=0, completed_requests=0, succeeded_requests=0, failed_requests=0, object='batch', metadata=None, model='mistral-ocr-latest', agent_id=None, output_file=None, error_file=None, outputs=None, started_at=None, completed_at=None)
Poll job until completion, printing only on status change
482ba9ee-559f-4b11-8f60-8aabe739ad05 CANCELLATION_REQUESTED 1770393138
40e30b98-dc12-412a-a5ce-52818f3a097a CANCELLATION_REQUESTED 1770392814
39e42cd1-b2e4-423a-8d63-ca17697efac8 CANCELLED 1770392339
23450f30-0d63-43be-83ce-206102e47c6d CANCELLATION_REQUESTED 1770390898
60e1ce21-6af8-4ad3-862d-6df52bd34cc7 CANCELLATION_REQUESTED 1770389326
ed556e64-884e-4349-9579-ed37af5e6920 CANCELLED 1770373917
2fa17c1d-959a-46ae-b0b2-7d966eb2c56a CANCELLED 1770371001
e7d64ad4-8394-4c9a-a3f4-d36a1681afcd CANCELLED 1770310312
9683cafb-a3a4-4346-934b-e6b06865657e CANCELLED 1770307850
944d4d93-271e-4994-8c72-dd1cf54d1182 CANCELLED 1770307352
('CANCELLED', [])
Download and parse batch job results
(1, dict_keys(['id', 'custom_id', 'response', 'error']))
Each result contains the custom ID you provided, the OCR response with pages and markdown content, and any errors. Images are embedded as base64 strings and need to be decoded before saving.
Upload PDFs and prepare batch entries
Submit batch, wait for completion, and download results
OCR results contain markdown for each page, plus any extracted images as base64-encoded data. The functions below save each page as a separate markdown file in a folder named after the PDF, with images in an img subfolder.
Adapt single OCR response to batch result format for uniform downstream processing
Save all images from a page to directory
Save single page markdown and images
For a PDF named paper.pdf, the output structure will be:
md/
paper/
page_1.md
page_2.md
img/
img-0.jpeg
img-1.jpeg
Each page is saved as a separate markdown file, making it easy to process or view individual pages.
Save markdown pages and images from OCR response to output directory
Path('files/test/md/attention-is-all-you-need')
img/ page_11.md page_14.md page_3.md page_6.md page_9.md
page_1.md page_12.md page_15.md page_4.md page_7.md
page_10.md page_13.md page_2.md page_5.md page_8.md
files/test/md:
attention-is-all-you-need/ resnet/
files/test/md/attention-is-all-you-need:
img/ page_11.md page_14.md page_3.md page_6.md page_9.md
page_1.md page_12.md page_15.md page_4.md page_7.md
page_10.md page_13.md page_2.md page_5.md page_8.md
files/test/md/attention-is-all-you-need/img:
img-0.jpeg img-2.jpeg img-4.jpeg img-6.jpeg
img-1.jpeg img-3.jpeg img-5.jpeg
files/test/md/resnet:
img/ page_10.md page_12.md page_3.md page_5.md page_7.md page_9.md
page_1.md page_11.md page_2.md page_4.md page_6.md page_8.md
files/test/md/resnet/img:
img-0.jpeg img-2.jpeg img-4.jpeg img-6.jpeg
img-1.jpeg img-3.jpeg img-5.jpeg
ocr_pdf is the main entry point. It auto-detects the best mode: single PDFs use direct OCR for immediate results; multiple PDFs use the batch API for cost savings.
Get list of PDFs from file or folder
For instance:
([Path('files/test/attention-is-all-you-need.pdf')],
[Path('files/test/attention-is-all-you-need.pdf'), Path('files/test/resnet.pdf')])
OCR PDF(s) and save results. Single PDF uses direct mode; multiple uses batch
Path('files/test/md_single/attention-is-all-you-need')
files/test/md_single/attention-is-all-you-need:
img page_11.md page_14.md page_3.md page_6.md page_9.md
page_1.md page_12.md page_15.md page_4.md page_7.md
page_10.md page_13.md page_2.md page_5.md page_8.md
files/test/md_single/attention-is-all-you-need/img:
img-0.jpeg img-2.jpeg img-4.jpeg img-6.jpeg
img-1.jpeg img-3.jpeg img-5.jpeg
files/test/md_all:
attention-is-all-you-need/ iom-hoa-short/ resnet/
files/test/md_all/attention-is-all-you-need:
img/ page_11.md page_14.md page_3.md page_6.md page_9.md
page_1.md page_12.md page_15.md page_4.md page_7.md
page_10.md page_13.md page_2.md page_5.md page_8.md
files/test/md_all/attention-is-all-you-need/img:
img-0.jpeg img-2.jpeg img-4.jpeg img-6.jpeg
img-1.jpeg img-3.jpeg img-5.jpeg
files/test/md_all/iom-hoa-short:
img/ page_14.md page_2.md page_25.md page_30.md page_8.md
page_1.md page_15.md page_20.md page_26.md page_31.md page_9.md
page_10.md page_16.md page_21.md page_27.md page_4.md
page_11.md page_17.md page_22.md page_28.md page_5.md
page_12.md page_18.md page_23.md page_29.md page_6.md
page_13.md page_19.md page_24.md page_3.md page_7.md
files/test/md_all/iom-hoa-short/img:
img-0.jpeg img-2.jpeg img-4.jpeg img-6.jpeg
img-1.jpeg img-3.jpeg img-5.jpeg
files/test/md_all/resnet:
img/ page_10.md page_12.md page_3.md page_5.md page_7.md page_9.md
page_1.md page_11.md page_2.md page_4.md page_6.md page_8.md
files/test/md_all/resnet/img:
img-0.jpeg img-10.jpeg img-2.jpeg img-4.jpeg img-6.jpeg img-8.jpeg
img-1.jpeg img-11.jpeg img-3.jpeg img-5.jpeg img-7.jpeg img-9.jpeg
Helper functions for working with OCR results and preparing PDFs for processing.
After running ocr_pdf, use read_pgs to load the markdown content. Set join=True (default) to get a single string suitable for LLM processing, or join=False to get a list of pages for individual page analysis.
Read specific page or all pages from OCR output directory
For instance to read all pages:
# Deep Residual Learning for Image Recognition
Kaiming He
Xiangyu Zhang
Shaoqing Ren
Jian Sun
Microsoft Research
{kahe, v-xiangz, v-shren, jiansun}@microsoft.com
# Abstract
Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreference
Or a list of markdown pages (that you can further index):
(12,
'# Deep Residual Learning for Image Recognition\n\nKaiming He\n\nXiangyu Zhang\n\nShaoqing Ren\n\nJian Sun\n\nMicrosoft Research\n\n{kahe, v-xiangz, v-shren, jiansun}@microsoft.com\n\n# Abstract\n\nDeeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreference')
If you need to OCR just a subset of pages (e.g., to test on a few pages before processing a large document), use subset_pdf to extract a page range first.
Extract page range from PDF and save with range suffix