PDF内容上下文分块中标题与子标题嵌套解析逻辑的复现问题
PDF内容上下文分块中标题与子标题嵌套解析逻辑的复现问题
嘿,我看你在折腾Google Docs转JSON后的标题-子标题嵌套解析,要做上下文分块对吧?你贴的代码已经搭了架子,但没写完,我来帮你补全完整逻辑,顺便把嵌套解析的核心思路掰扯清楚——这种层级化的标题处理,栈结构是最靠谱的工具,我之前处理过好几个类似的需求,亲测好用。
先给你补全并优化后的完整代码:
def format(json_data): """ Extracts document title, headings, subheadings (if present), and content in a specific JSON format. Args: json_data: A dictionary containing the parsed JSON data of the Google Doc. Returns: A list containing a dictionary with the document title and another dictionary for each heading with its subheading (if any) and content. """ extracted_data = [] current_heading = None current_subheading = None current_heading_level = None # 用栈管理标题层级,栈元素存(层级数字, 对应标题/子标题的字典) stack = [] # 辅助函数:提取段落完整文本(处理带格式的多元素段落) def get_paragraph_text(paragraph): text = "" for elem in paragraph["elements"]: if "textRun" in elem: text += elem["textRun"]["content"] return text.strip() for element in json_data["body"]["content"]: if "paragraph" in element: paragraph = element["paragraph"] paragraph_style = paragraph.get("paragraphStyle", {}) element_type = paragraph_style.get("namedStyleType") if element_type == "TITLE": # 处理文档主标题 title_text = get_paragraph_text(paragraph) extracted_data.append({"document_title": title_text}) elif element_type in ["HEADING_1", "HEADING_2", "HEADING_3"]: # 把标题类型映射成数字层级,方便比较嵌套关系 heading_level_map = {"HEADING_1": 1, "HEADING_2": 2, "HEADING_3": 3} current_level = heading_level_map[element_type] heading_text = get_paragraph_text(paragraph) # 核心嵌套处理:弹出栈中所有层级>=当前层级的元素,回到父级层级 while stack and stack[-1][0] >= current_level: stack.pop() # 按层级分别处理 if current_level == 1: # 一级标题,作为独立块加入结果 current_heading = {"heading": heading_text, "subheadings": []} extracted_data.append(current_heading) stack.append((current_level, current_heading)) current_subheading = None elif current_level == 2: # 二级标题,挂到最近的一级标题下 if stack: parent_heading = stack[-1][1] current_subheading = {"subheading": heading_text, "content": []} parent_heading["subheadings"].append(current_subheading) stack.append((current_level, current_subheading)) elif current_level == 3: # 三级标题,挂到最近的二级标题下 if stack: parent_subheading = stack[-1][1] current_sub_subheading = {"sub_subheading": heading_text, "content": []} parent_subheading["content"].append(current_sub_subheading) stack.append((current_level, current_sub_subheading)) current_heading_level = current_level elif element_type == "NORMAL_TEXT": # 处理普通内容,追加到最近的标题/子标题块中 content_text = get_paragraph_text(paragraph) if not content_text: continue # 跳过空段落 if current_subheading: current_subheading["content"].append(content_text) elif current_heading: if "content" not in current_heading: current_heading["content"] = [] current_heading["content"].append(content_text) return extracted_data
再给你唠唠核心逻辑:
- 用
get_paragraph_text是因为Google Docs的段落经常会拆成多个元素(比如一段里混了粗体和普通文本),只取第一个元素会丢内容; - 栈的用法是关键:标题嵌套本质就是栈结构——进入子层级就压栈,回到父层级就弹栈,绝对不会搞混层级关系;
- 层级映射成数字后,判断嵌套关系超级直观:比如遇到H2,就把栈里所有层级>=2的元素弹掉,剩下的栈顶就是它的父H1;
- 容错处理也考虑到了:空段落直接跳过,避免把无效内容加入分块里。
如果要扩展支持更多层级(比如H4/H5),只需要在heading_level_map里加对应映射,再补充对应的层级判断逻辑就行。
备注:内容来源于stack exchange,提问作者whysumedh
相关产品推荐
相关产品推荐

