CoreSectionsOutput
def CoreSectionsOutput(
**data:Any
)->None:Identify the core sections of the report
Evaluation reports run from 50 to over 200 pages. The mapping against the SRF and GCM uses only the core sections of a report, such as the executive summary, introduction, conclusions and recommendations. Leaving out the rest removes tangential content. It also stops a theme the report only mentions in passing from counting as relevant.
There are two ways to choose the core sections. A person selects their headings in the curator app, and extract_selected extracts them. Or extract_sections asks an LLM to choose them from the report’s table of contents.
Identify the core sections of the report
Parse the headings of md with exhash’s open_doc into a nested dict keyed by heading title
Nested dict of section titles; each value holds its section’s Markdown, heading and subsections included, in text
create_heading_dict turns a report’s Markdown into a nested dict keyed by heading title. Each value holds its section’s Markdown in text, with its heading and subsections. Two headings with the same title under the same parent share one key, and the dict keeps the last of them.
To try it on a real report, load its OCR’d pages with read_pgs:
img page_16.md page_23.md page_30.md page_38.md page_45.md
page_1.md page_17.md page_24.md page_31.md page_39.md page_46.md
page_10.md page_18.md page_25.md page_32.md page_4.md page_47.md
page_11.md page_19.md page_26.md page_33.md page_40.md page_5.md
page_12.md page_2.md page_27.md page_34.md page_41.md page_6.md
page_13.md page_20.md page_28.md page_35.md page_42.md page_7.md
page_14.md page_21.md page_29.md page_36.md page_43.md page_8.md
page_15.md page_22.md page_3.md page_37.md page_44.md page_9.md
The examples below use this short report:
sample_md = """# Report Title ... page 1
## Executive Summary ... page 1
This is a summary of key findings.
## 1. Introduction ... page 2
Background information here.
### 1.1 Objectives ... page 2
The objectives are...
## 2. Findings ... page 3
Detailed findings.
## 3. Conclusions ... page 5
Main conclusions.
## 4. Recommendations ... page 6
Key recommendations."""{'Report Title ... page 1': {'Executive Summary ... page 1': {},
'1. Introduction ... page 2': {'1.1 Objectives ... page 2': {}},
'2. Findings ... page 3': {},
'3. Conclusions ... page 5': {},
'4. Recommendations ... page 6': {}}}
['# Report Title ... page 1',
'## Executive Summary ... page 1',
'## 1. Introduction ... page 2',
'### 1.1 Objectives ... page 2',
'## 2. Findings ... page 3',
'## 3. Conclusions ... page 5',
'## 4. Recommendations ... page 6']
The curator app saves a selection as raw headings, such as "## 1. Introduction ... page 4". heading_to_key strips the # prefix to get the title. find_path then finds the titles that lead from the top of the dict to that heading.
Title of the raw heading hdg, without its # prefix
Find the path of titles to the heading key in hdgs
sample_hdgs = create_heading_dict(sample_md)
assert find_path(sample_hdgs, "1. Introduction ... page 2") == [
"Report Title ... page 1",
"1. Introduction ... page 2"
]
assert find_path(sample_hdgs, "1.1 Objectives ... page 2") == [
"Report Title ... page 1",
"1. Introduction ... page 2",
"1.1 Objectives ... page 2"
]
assert find_path(sample_hdgs, "nonexistent") == []A selection can hold a section and one of its subsections, such as 1. Introduction and 1.1 Objectives. Extracting both would repeat the subsection’s text. rm_nested drops every path that starts with another path in the list. It returns the paths shortest first, and the extracted text follows that order.
Drop the paths that extend another path in paths
[['Report Title ... page 1', '1. Introduction ... page 2'],
['Report Title ... page 1', '3. Conclusions ... page 5']]
Find the paths to selected_headings, skipping headings that are not in hdgs
A subsection’s heading gives the full path, from the report title down:
get_text follows a path of titles down the dict and returns that section’s Markdown, subsections included.
Get the Markdown of the section at path ks in hdgs
identify_core_sections sends the table of contents to an LLM, with the select_sections prompt. The table of contents is the nested dict of titles, without the sections’ text. The LLM replies with section_paths, the paths to the core sections, and its reasoning. The LLM handles titles in other languages, varied naming and unusual structures.
async def identify_core_sections(
hdgs:dict, # Nested dict from `create_heading_dict`
sp:str=None, # System prompt; `None` uses the `select_sections` prompt
response_format:type[pydantic.main.BaseModel]=CoreSectionsOutput, # Pydantic model the reply must follow
model:str='claude-sonnet-4-5', # LLM to use
)->dict: # `section_paths` and `reasoning`Ask an LLM for the paths to the core sections of the report outlined by hdgs
{'section_paths': [['Report Title ... page 1', 'Executive Summary ... page 1'],
['Report Title ... page 1', '1. Introduction ... page 2'],
['Report Title ... page 1',
'1. Introduction ... page 2',
'1.1 Objectives ... page 2'],
['Report Title ... page 1', '3. Conclusions ... page 5'],
['Report Title ... page 1', '4. Recommendations ... page 6']],
'reasoning': "Selected all five key sections that reveal core themes: Executive Summary (page 1) provides high-level overview of what authors consider important; Introduction and Objectives (pages 2) establish the evaluation's purpose and scope; Conclusions (page 5) synthesize key findings and judgments; Recommendations (page 6) translate conclusions into actionable guidance. These sections total approximately 6 pages and represent the most strategic content for determining core themes. Excluded Findings section as it likely contains detailed evidence that supports rather than defines the core themes, though the selected sections should provide sufficient context about what findings matter most."}
extract_selected extracts the sections under a list of raw headings, such as a selection saved by the curator app. It needs no LLM. extract_sections does the same when you pass selected_headings. Without them, it asks identify_core_sections to choose the sections, and it is async for that reason. Both return the report’s first top-level title as an H1, followed by the text of each section. join_sections skips any path that is not in the report and logs a warning.
Get the report title from hdgs
Concatenate the report title and the text of each section in paths
Extract the sections under selected_headings, such as those chosen in the curator app
async def extract_sections(
md:str, # Markdown text of the full report
selected_headings:list[str]=None, # Raw headings like `"## 1. Intro ..."`; `None` asks the LLM
*, sp:str=None, # System prompt; `None` uses the `select_sections` prompt
response_format:type[pydantic.main.BaseModel]=CoreSectionsOutput, # Pydantic model the reply must follow
model:str='claude-sonnet-4-5', # LLM to use
)->str: # Report title and the text of each core sectionExtract the core sections of md, chosen by an LLM unless you pass selected_headings
Selecting 1. Introduction extracts its subsection too, under the report title:
# Report Title ... page 1
## 1. Introduction ... page 2
Background information here.
### 1.1 Objectives ... page 2
The objectives are...
An empty selection extracts nothing: