Extract

Parse a report’s headings, choose its core sections from a list of headings or with an LLM, and extract their text

Evaluation reports run from 50 to over 200 pages. The mapping against the SRF and GCM uses only the core sections of a report, such as the executive summary, introduction, conclusions and recommendations. Leaving out the rest removes tangential content. It also stops a theme the report only mentions in passing from counting as relevant.

There are two ways to choose the core sections. A person selects their headings in the curator app, and extract_selected extracts them. Or extract_sections asks an LLM to choose them from the report’s table of contents.


source

CoreSectionsOutput

def CoreSectionsOutput(
    **data:Any
)->None:

Identify the core sections of the report

Exported source
class CoreSectionsOutput(BaseModel):
    "Identify the core sections of the report"
    section_paths: list[list[str]]
    reasoning: str

source

create_heading_dict

def create_heading_dict(
    md:str, # Markdown text of full report
)->__main__.HeadingDict: # Nested dict of the report's headings keyed by title

Parse the headings of md with exhash’s open_doc into a nested dict keyed by heading title


source

HeadingDict

def HeadingDict(
    *args, **kwargs
):

Nested dict of section titles; each value holds its section’s Markdown, heading and subsections included, in text

create_heading_dict turns a report’s Markdown into a nested dict keyed by heading title. Each value holds its section’s Markdown in text, with its heading and subsections. Two headings with the same title under the same parent share one key, and the dict keeps the last of them.

To try it on a real report, load its OCR’d pages with read_pgs:

from mistocr.core import read_pgs
!ls ../../pipeline/data/md/4341695461234eee3deb51ac68871109
img     page_16.md  page_23.md  page_30.md  page_38.md  page_45.md
page_1.md   page_17.md  page_24.md  page_31.md  page_39.md  page_46.md
page_10.md  page_18.md  page_25.md  page_32.md  page_4.md   page_47.md
page_11.md  page_19.md  page_26.md  page_33.md  page_40.md  page_5.md
page_12.md  page_2.md   page_27.md  page_34.md  page_41.md  page_6.md
page_13.md  page_20.md  page_28.md  page_35.md  page_42.md  page_7.md
page_14.md  page_21.md  page_29.md  page_36.md  page_43.md  page_8.md
page_15.md  page_22.md  page_3.md   page_37.md  page_44.md  page_9.md
fname = "../../pipeline/data/md/4341695461234eee3deb51ac68871109"
sample_md = read_pgs(fname, join=True)

The examples below use this short report:

sample_md = """# Report Title ... page 1

## Executive Summary ... page 1

This is a summary of key findings.

## 1. Introduction ... page 2

Background information here.

### 1.1 Objectives ... page 2

The objectives are...

## 2. Findings ... page 3

Detailed findings.

## 3. Conclusions ... page 5

Main conclusions.

## 4. Recommendations ... page 6

Key recommendations."""
hdgs = create_heading_dict(sample_md)
hdgs
{'Report Title ... page 1': {'Executive Summary ... page 1': {},
  '1. Introduction ... page 2': {'1.1 Objectives ... page 2': {}},
  '2. Findings ... page 3': {},
  '3. Conclusions ... page 5': {},
  '4. Recommendations ... page 6': {}}}
from mistocr.refine import fmt_hdgs_idx
print(fmt_hdgs_idx(hdgs))
0. Report Title ... page 1
re.findall(r'^#+\s+.*$', sample_md, re.MULTILINE)
['# Report Title ... page 1',
 '## Executive Summary ... page 1',
 '## 1. Introduction ... page 2',
 '### 1.1 Objectives ... page 2',
 '## 2. Findings ... page 3',
 '## 3. Conclusions ... page 5',
 '## 4. Recommendations ... page 6']

From Raw Headings to Dictionary Paths

The curator app saves a selection as raw headings, such as "## 1. Introduction ... page 4". heading_to_key strips the # prefix to get the title. find_path then finds the titles that lead from the top of the dict to that heading.


source

heading_to_key

def heading_to_key(
    hdg:str, # Raw markdown heading like "## 1. Intro ..."
)->str: # Heading text without markdown prefix

Title of the raw heading hdg, without its # prefix

assert heading_to_key("## 1. Introduction ... page 4") == "1. Introduction ... page 4"
assert heading_to_key("### 2.1. Context ... page 5") == "2.1. Context ... page 5"
assert heading_to_key("# Title") == "Title"

source

find_path

def find_path(
    hdgs:dict, # Nested dict from `create_heading_dict`
    key:str, # Heading title to find
)->list[str]: # Titles from the top level down to `key`; empty if `key` is not in `hdgs`

Find the path of titles to the heading key in hdgs

sample_hdgs = create_heading_dict(sample_md)
assert find_path(sample_hdgs, "1. Introduction ... page 2") == [
    "Report Title ... page 1",
    "1. Introduction ... page 2"
]
assert find_path(sample_hdgs, "1.1 Objectives ... page 2") == [
    "Report Title ... page 1",
    "1. Introduction ... page 2",
    "1.1 Objectives ... page 2"
]
assert find_path(sample_hdgs, "nonexistent") == []

A selection can hold a section and one of its subsections, such as 1. Introduction and 1.1 Objectives. Extracting both would repeat the subsection’s text. rm_nested drops every path that starts with another path in the list. It returns the paths shortest first, and the extracted text follows that order.


source

rm_nested

def rm_nested(
    paths:list[list[str]], # Section paths, each a list of titles
)->list[list[str]]: # The paths that don't extend another path, shortest first

Drop the paths that extend another path in paths

nested_paths = [
    ['Report Title ... page 1', '1. Introduction ... page 2'],
    ['Report Title ... page 1', '1. Introduction ... page 2', '1.1 Objectives ... page 2'],
    ['Report Title ... page 1', '3. Conclusions ... page 5']
]

rm_nested(nested_paths)
[['Report Title ... page 1', '1. Introduction ... page 2'],
 ['Report Title ... page 1', '3. Conclusions ... page 5']]

source

headings_to_paths

def headings_to_paths(
    hdgs:dict, # Nested dict from `create_heading_dict`
    selected_headings:list[str], # Raw headings like `"## 1. Intro ..."`
)->list[list[str]]: # Paths to the selected sections, without nested ones

Find the paths to selected_headings, skipping headings that are not in hdgs

A subsection’s heading gives the full path, from the report title down:

selected = ["### 1.1 Objectives ... page 2"]
paths = headings_to_paths(sample_hdgs, selected)
assert paths == [[
    "Report Title ... page 1",
    "1. Introduction ... page 2", 
    "1.1 Objectives ... page 2"
]]

Extracting a section

get_text follows a path of titles down the dict and returns that section’s Markdown, subsections included.


source

get_text

def get_text(
    ks:list[str], # Titles from the top level down to the section
    hdgs:dict, # Nested dict from `create_heading_dict`
)->str: # The section's Markdown, subsections included

Get the Markdown of the section at path ks in hdgs

path = ['Report Title ... page 1', '3. Conclusions ... page 5']
print(get_text(path, hdgs))
## 3. Conclusions ... page 5

Main conclusions.

Choosing sections with an LLM

identify_core_sections sends the table of contents to an LLM, with the select_sections prompt. The table of contents is the nested dict of titles, without the sections’ text. The LLM replies with section_paths, the paths to the core sections, and its reasoning. The LLM handles titles in other languages, varied naming and unusual structures.


source

identify_core_sections

async def identify_core_sections(
    hdgs:dict, # Nested dict from `create_heading_dict`
    sp:str=None, # System prompt; `None` uses the `select_sections` prompt
    response_format:type[pydantic.main.BaseModel]=CoreSectionsOutput, # Pydantic model the reply must follow
    model:str='claude-sonnet-4-5', # LLM to use
)->dict: # `section_paths` and `reasoning`

Ask an LLM for the paths to the core sections of the report outlined by hdgs

sections = await identify_core_sections(hdgs)
sections
{'section_paths': [['Report Title ... page 1', 'Executive Summary ... page 1'],
  ['Report Title ... page 1', '1. Introduction ... page 2'],
  ['Report Title ... page 1',
   '1. Introduction ... page 2',
   '1.1 Objectives ... page 2'],
  ['Report Title ... page 1', '3. Conclusions ... page 5'],
  ['Report Title ... page 1', '4. Recommendations ... page 6']],
 'reasoning': "Selected all five key sections that reveal core themes: Executive Summary (page 1) provides high-level overview of what authors consider important; Introduction and Objectives (pages 2) establish the evaluation's purpose and scope; Conclusions (page 5) synthesize key findings and judgments; Recommendations (page 6) translate conclusions into actionable guidance. These sections total approximately 6 pages and represent the most strategic content for determining core themes. Excluded Findings section as it likely contains detailed evidence that supports rather than defines the core themes, though the selected sections should provide sufficient context about what findings matter most."}

Extracting the core sections

extract_selected extracts the sections under a list of raw headings, such as a selection saved by the curator app. It needs no LLM. extract_sections does the same when you pass selected_headings. Without them, it asks identify_core_sections to choose the sections, and it is async for that reason. Both return the report’s first top-level title as an H1, followed by the text of each section. join_sections skips any path that is not in the report and logs a warning.


source

get_h1_title

def get_h1_title(
    hdgs:dict, # Nested dict from `create_heading_dict`
)->str: # The first top-level title as a Markdown H1; empty if there are no headings

Get the report title from hdgs


source

join_sections

def join_sections(
    hdgs:dict, # Nested dict from `create_heading_dict`
    paths:list[list[str]], # Paths to the sections to extract
)->str: # Report title followed by each section's text; empty if no paths

Concatenate the report title and the text of each section in paths


source

extract_selected

def extract_selected(
    md:str, # Markdown text of the full report
    selected_headings:list[str], # Raw headings like `"## 1. Intro ..."`
)->str: # Report title and the text of each selected section

Extract the sections under selected_headings, such as those chosen in the curator app


source

extract_sections

async def extract_sections(
    md:str, # Markdown text of the full report
    selected_headings:list[str]=None, # Raw headings like `"## 1. Intro ..."`; `None` asks the LLM
    *, sp:str=None, # System prompt; `None` uses the `select_sections` prompt
    response_format:type[pydantic.main.BaseModel]=CoreSectionsOutput, # Pydantic model the reply must follow
    model:str='claude-sonnet-4-5', # LLM to use
)->str: # Report title and the text of each core section

Extract the core sections of md, chosen by an LLM unless you pass selected_headings

Selecting 1. Introduction extracts its subsection too, under the report title:

text = extract_selected(sample_md, ["## 1. Introduction ... page 2"])
assert "# Report Title ... page 1" in text
assert "Background information here" in text
assert "1.1 Objectives" in text
print(text)
# Report Title ... page 1
## 1. Introduction ... page 2

Background information here.

### 1.1 Objectives ... page 2

The objectives are...

An empty selection extracts nothing:

text = extract_selected(sample_md, [])
assert text == ""
text
''