IOMEval
IOMEval streamlines the mapping of IOM evaluation reports against strategic frameworks like the Strategic Results Framework (SRF) and the Global Compact for Migration (GCM). It uses LLMs to process PDF reports, extract key sections, and tag (match) them to framework components, turning dispersed, untagged evaluation documents into structured evidence maps that can be searched by framework components (for example, finding all evaluation findings mapped to a specific GCM objective).
The Challenge Addressed
UN agencies produce a large number of evaluation reports. For IOM, this body of knowledge is extensive and variegated, but putting it to practical use becomes more challenging as volume increases. Critically, the metadata of IOM evaluation reports does not indicate which elements in the IOM Strategic Results Framework (SRF), or in the Global Compact for Migration (GCM), are addressed by the evaluation. This is a major gap that limits the ability to connect evaluation evidence with the two key strategic frameworks of the organization.
Manual tagging of evaluation reports against the IOM SRF and the GCM is extremely challenging due to the limited resources that IOM has at its disposal for evaluation in general. Time constraints of IOM evaluators and other staff are also exacerbated by the shrinkage of the organization’s budget in the context of the broader “humanitarian reset”. In addition to this, tagging IOM evaluation reports against SRF elements is cognitively taxing due to the sheer number of elements in these frameworks (the GCM has 23 objectives; the SRF has more than one hundred outputs).
What This Enables
Addressing the “tagging” challenge enables the creation of evidence maps (visual tools that systematically display what evaluation and research exist for specific topics, and where evidence may be missing) that would otherwise not have been possible to produce. Maps in turn help answer questions like: Which framework elements are well-covered by existing evaluations? Where are the knowledge gaps that should prioritize future evaluation work? Which themes have enough evidence for a dedicated synthesis report?
Key Features
- Automated PDF Processing: Download and OCR evaluation reports
- Intelligent Section Extraction: LLM-powered extraction of executive summaries, findings, conclusions, and recommendations
- Strategic Framework Mapping: Map report content against the IOM Strategic Results Framework (outputs, enablers and cross-cutting priorities) and the Global Compact for Migration (objectives)
- Checkpoint/Resume: Stop processing at any time and pick up where you left off without losing progress
- Granular Control: Run the entire processing pipeline for a report in one command, or execute individual steps (PDF extraction, section identification, specific framework theme mapping) separately for more flexibility
Installation
Install from PyPI:
pip install iomevalOr install the latest development version from GitHub:
pip install git+https://github.com/franckalbinet/iomeval.gitConfiguration
Core Dependencies
iomeval relies on two key libraries:
API Keys
iomeval automatically loads API keys on import. You have two options:
Option 1: Environment variables (recommended for production)
export ANTHROPIC_API_KEY='your-key-here'
export MISTRAL_API_KEY='your-key-here'Option 2: .env file (convenient for development)
Create a .env file in your project root:
ANTHROPIC_API_KEY=your-key-here
MISTRAL_API_KEY=your-key-here
Since fastllm supports several LLM providers, you can configure other providers (OpenAI, Google, etc.) by setting their respective API keys using either method.
Quick Start
First, prepare your evaluation report metadata. Export the CSV file from the IOM Evaluation Repository containing report records and their associated PDF links, then convert to JSON:
from iomeval.readers import IOMRepoReader
reader = IOMRepoReader('evaluation-search-export.csv')
reader.to_json('evaluations.json')Now process an evaluation report through the complete pipeline:
from iomeval.readers import load_evals
from iomeval.pipeline import run_pipeline
evals = load_evals('evaluations.json')
url = "https://evaluation.iom.int/sites/g/files/tmzbdl151/files/docs/resources/Abridged%20Evaluation%20Report_%20Final_Olta%20NDOJA.pdf"
result = await run_pipeline(evals, url=url, base_path='data', add_img_desc=False, model='claude-haiku-4-5')
result.statusThe pipeline downloads the PDF, runs OCR, then stops with status 'awaiting_curation': select the report’s core sections in the curator app (06_curator.ipynb), then run the same call again. It then maps the selected sections against each strategic framework in turn (SRF Enablers, SRF Cross-cutting Priorities, GCM Objectives, and SRF Outputs), and returns status 'completed', with the scores in result.report.mappings.
State is saved under base_path after each stage, so a re-run resumes where it stopped. batch_run does the same for many reports.
The prompts used for extraction and framework mapping are available in the prompts directory.
Curating Reports
Before mapping, a person reviews each OCR’d report in the curator app, a small local web app. It works in two steps:
- Clean headings. OCR can get heading text or levels wrong. You correct them in a form, and saving rewrites the report’s Markdown pages.
- Select sections. You tick the sections to map, usually the executive summary, introduction, conclusions and recommendations. A counter shows their size against a 15,000-token budget.
The selection is saved with the report’s results, with status 'sections_selected'. The next run_pipeline call maps those sections.
Why a person does this: OCR headings need checking before they can be trusted, and mapping only the core sections costs less and keeps scores focused on the report’s findings.
Why it runs locally: the app edits the OCR’d Markdown files in place, so it runs on the machine that holds the reports. It reads them from BASE_PATH in iomeval.curator, which is ../data relative to the folder you start it from. That folder has the same layout as the base_path given to run_pipeline, with md/ and results/ inside.
To serve it, start it from a folder next to your data folder:
python -m iomeval.curator --port 5001Then open http://localhost:5001. You can also serve it from 06_curator.ipynb with srv = JupyUvi(app).
Detailed Workflow
For more control over individual pipeline stages, see the module documentation:
- Loading evaluation metadata: See readers for working with IOM evaluation data
- Downloading and OCR: See downloaders and core for PDF processing
- Section extraction: See extract for extracting executive summaries, findings, conclusions, and recommendations
- Framework mapping: See mapper for mapping to SRF enablers, cross-cutting priorities, GCM objectives, and SRF outputs
- Pipeline control: See pipeline for granular control over the full pipeline and checkpoint/resume functionality
Development
iomeval is built with nbdev, which means the entire library is developed in Jupyter notebooks. The notebooks serve as both documentation and source code.
Setup for development
git clone https://github.com/franckalbinet/iomeval.git
cd iomeval
pip install -e '.[dev]'Key nbdev commands
nbdev-test # Run tests in notebooks
nbdev-export # Export notebooks to Python modules
nbdev-preview # Preview documentation site
nbdev-prepare # Export, test, clean notebooks, and render README (run before committing)Workflow
- Make changes in the
.ipynbnotebook files - Run
nbdev-prepareto export code and run tests - Commit both notebooks and exported Python files
- Documentation is automatically generated from the notebooks
Learn more about nbdev’s literate programming approach in the nbdev documentation.
Contributing
Contributions are welcome! Please: - Follow the existing notebook structure - Run nbdev-prepare before submitting PRs
