You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中将带层级标题的Markdown解析为JSON?

Markdown转层级化JSON解析问题

需求说明

我有大量包含多级标题的Markdown文件,需要将其解析为能区分各级标题对应文本及下属子标题的JSON结构。核心要求是标题层级正确嵌套:子标题归属于上级标题,同级标题处于同一层级。示例如下:

输入Markdown

outer1
outer2

# title 1
text1.1

## title 1.1
text1.1.1

# title 2
text 2.1

目标JSON

{
  "text": [
    "outer1",
    "outer2"
  ],
  "inner": [
    {
      "section": [
        {
          "title": "title 1",
          "inner": [
            {
              "text": [
                "text1.1"
              ],
              "inner": [
                {
                  "section": [
                    {
                      "title": "title 1.1",
                      "inner": [
                        {
                          "text": [
                            "text1.1.1"
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "title": "title 2",
          "inner": [
            {
              "text": [
                "text2.1"
              ]
            }
          ]
        }
      ]
    }
  ]
}

当前尝试与问题

我用pyparsing尝试实现,但无法处理标题层级的计数逻辑(判断新标题的#数量是否小于等于当前层级),导致同级标题无法正确区分。当前代码如下:

import pyparsing as pp
from pytest import test_eq

section = pp.Forward()("section")
inner_block = pp.Forward()("inner")

start_section = pp.OneOrMore(pp.Word("#"))
title_section = line
title = start_section.suppress() + title_section('title')

line = pp.Combine(
    pp.OneOrMore(pp.Word(pp.unicode.Latin1.printables), stop_on=pp.LineEnd()),
    join_string=' ', adjacent=False)
text = ~title + pp.OneOrMore(line, stop_on=(pp.LineEnd() + pp.FollowedBy("#")))

inner_block << pp.Group(section | (text('text') + pp.Optional(section.set_parse_action(foo))))

section << pp.Group(title + pp.Optional(inner_block))

markdown = pp.OneOrMore(inner_block)


test = """\
out1
out2

# title 1
text1.1

# title 2
text2.1

"""

res = markdown.parse_string(test, parse_all=True).as_dict()
test_eq(res, dict(
    inner=[
        dict(
            text = ["out1", "out2"],
            section=[
                dict(title="title 1", inner=[
                    dict(
                        text=["text1.1"]
                    ),
                ]),
                dict(title="title 2", inner=[
                    dict(
                        text=["text2.1"]
                    ),
                ]),
            ]
        )
    ]
))

问题分析与解决方案

1. pyparsing的处理思路

pyparsing是上下文无关文法解析器,默认不跟踪解析状态,但可以通过**解析动作(parse action)**手动维护层级上下文:

  • 解析标题时,记录当前标题的层级(#的数量)
  • 用栈结构保存当前层级的节点,遇到新标题时,对比层级调整栈的深度:层级更高则压入新节点,层级更低则弹出栈顶直到匹配目标层级,再添加同级节点
  • 解析文本时,将内容添加到栈顶节点的文本列表中

2. 替代解析器推荐

如果不想手动维护状态,可选择自带AST(抽象语法树)生成的Markdown解析库:

  • Python-Markdown:官方解析库,会将文档转为HTML DOM结构,通过遍历DOM即可提取层级关系,自动处理标题嵌套
  • mistune:轻量快速的Markdown解析器,支持生成AST,遍历AST可直接构建目标层级结构

3. 纯Python实现思路

轻量方案可采用逐行处理+栈维护层级:

  • 初始化栈,栈中元素为当前层级的节点(包含title、text列表、inner子节点)
  • 逐行读取Markdown:
    • 识别标题行:统计#数量得到层级,提取标题文本
    • 调整栈深度:新层级大于栈深度则压入新节点;小于则弹出栈直到栈深度等于新层级-1,再添加同级节点
    • 识别文本行:将内容添加到栈顶节点的text列表中
  • 最终将栈顶节点整理为目标JSON结构

示例:用Python-Markdown实现

import markdown
import json
from markdown.treeprocessors import Treeprocessor
from markdown.extensions import Extension

class HierarchyCollector(Treeprocessor):
    def __init__(self):
        self.result = {"text": [], "inner": []}
        self.stack = [self.result]
    
    def run(self, root):
        for elem in root:
            if elem.tag.startswith("h") and elem.tag[1].isdigit():
                level = int(elem.tag[1])
                title = elem.text.strip() if elem.text else ""
                # 调整栈到目标层级的父节点
                while len(self.stack) > level:
                    self.stack.pop()
                # 创建新的section节点
                section_node = {"title": title, "inner": []}
                # 确保当前栈顶有section容器
                if not any(item.get("section") for item in self.stack[-1]["inner"]):
                    self.stack[-1]["inner"].append({"section": []})
                # 添加到section列表
                self.stack[-1]["inner"][0]["section"].append(section_node)
                self.stack.append(section_node)
            elif elem.tag == "p":
                text = elem.text.strip() if elem.text else ""
                if text:
                    # 确保当前栈顶有text列表
                    if "text" not in self.stack[-1]:
                        self.stack[-1]["text"] = []
                    self.stack[-1]["text"].append(text)
        return root

class HierarchyExtension(Extension):
    def extendMarkdown(self, md):
        md.treeprocessors.register(HierarchyCollector(), 'hierarchy_collector', 100)

# 测试解析
md_text = """
outer1
outer2

# title 1
text1.1

## title 1.1
text1.1.1

# title 2
text 2.1
"""

md = markdown.Markdown(extensions=[HierarchyExtension()])
md.convert(md_text)
result = md.treeprocessors['hierarchy_collector'].result
print(json.dumps(result, indent=2))

运行后会输出符合要求的JSON结构。

内容的提问来源于stack exchange,提问作者matanc1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 05:10:51