You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF内容上下文分块中标题与子标题嵌套解析逻辑的复现问题

PDF内容上下文分块中标题与子标题嵌套解析逻辑的复现问题

嘿,我看你在折腾Google Docs转JSON后的标题-子标题嵌套解析,要做上下文分块对吧?你贴的代码已经搭了架子,但没写完,我来帮你补全完整逻辑,顺便把嵌套解析的核心思路掰扯清楚——这种层级化的标题处理,栈结构是最靠谱的工具,我之前处理过好几个类似的需求,亲测好用。

先给你补全并优化后的完整代码:

def format(json_data):
    """
    Extracts document title, headings, subheadings (if present), and content in a specific JSON format.

    Args:
        json_data: A dictionary containing the parsed JSON data of the Google Doc.

    Returns:
        A list containing a dictionary with the document title and another dictionary
        for each heading with its subheading (if any) and content.
    """
    extracted_data = []
    current_heading = None
    current_subheading = None
    current_heading_level = None
    # 用栈管理标题层级,栈元素存(层级数字, 对应标题/子标题的字典)
    stack = []

    # 辅助函数:提取段落完整文本(处理带格式的多元素段落)
    def get_paragraph_text(paragraph):
        text = ""
        for elem in paragraph["elements"]:
            if "textRun" in elem:
                text += elem["textRun"]["content"]
        return text.strip()

    for element in json_data["body"]["content"]:
        if "paragraph" in element:
            paragraph = element["paragraph"]
            paragraph_style = paragraph.get("paragraphStyle", {})
            element_type = paragraph_style.get("namedStyleType")

            if element_type == "TITLE":
                # 处理文档主标题
                title_text = get_paragraph_text(paragraph)
                extracted_data.append({"document_title": title_text})
            elif element_type in ["HEADING_1", "HEADING_2", "HEADING_3"]:
                # 把标题类型映射成数字层级,方便比较嵌套关系
                heading_level_map = {"HEADING_1": 1, "HEADING_2": 2, "HEADING_3": 3}
                current_level = heading_level_map[element_type]
                heading_text = get_paragraph_text(paragraph)

                # 核心嵌套处理:弹出栈中所有层级>=当前层级的元素,回到父级层级
                while stack and stack[-1][0] >= current_level:
                    stack.pop()

                # 按层级分别处理
                if current_level == 1:
                    # 一级标题,作为独立块加入结果
                    current_heading = {"heading": heading_text, "subheadings": []}
                    extracted_data.append(current_heading)
                    stack.append((current_level, current_heading))
                    current_subheading = None
                elif current_level == 2:
                    # 二级标题,挂到最近的一级标题下
                    if stack:
                        parent_heading = stack[-1][1]
                        current_subheading = {"subheading": heading_text, "content": []}
                        parent_heading["subheadings"].append(current_subheading)
                        stack.append((current_level, current_subheading))
                elif current_level == 3:
                    # 三级标题,挂到最近的二级标题下
                    if stack:
                        parent_subheading = stack[-1][1]
                        current_sub_subheading = {"sub_subheading": heading_text, "content": []}
                        parent_subheading["content"].append(current_sub_subheading)
                        stack.append((current_level, current_sub_subheading))
                current_heading_level = current_level
            elif element_type == "NORMAL_TEXT":
                # 处理普通内容,追加到最近的标题/子标题块中
                content_text = get_paragraph_text(paragraph)
                if not content_text:
                    continue  # 跳过空段落
                if current_subheading:
                    current_subheading["content"].append(content_text)
                elif current_heading:
                    if "content" not in current_heading:
                        current_heading["content"] = []
                    current_heading["content"].append(content_text)
    return extracted_data

再给你唠唠核心逻辑:

  • 用get_paragraph_text是因为Google Docs的段落经常会拆成多个元素(比如一段里混了粗体和普通文本),只取第一个元素会丢内容;
  • 栈的用法是关键:标题嵌套本质就是栈结构——进入子层级就压栈,回到父层级就弹栈,绝对不会搞混层级关系;
  • 层级映射成数字后,判断嵌套关系超级直观:比如遇到H2,就把栈里所有层级>=2的元素弹掉,剩下的栈顶就是它的父H1;
  • 容错处理也考虑到了:空段落直接跳过,避免把无效内容加入分块里。

如果要扩展支持更多层级(比如H4/H5),只需要在heading_level_map里加对应映射,再补充对应的层级判断逻辑就行。

备注:内容来源于stack exchange,提问作者whysumedh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.17 08:48:13