You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取Markdown文件中各标题间的文本内容?

当然可以实现,下面给你两种实用的方法:

方法一:Python 脚本提取

用正则匹配三级标题,将内容按标题分割后提取每个标题对应的文本块,能自动处理空内容的情况:

import re

# 读取目标Markdown文件内容
with open("your_file.md", "r", encoding="utf-8") as f:
    md_content = f.read()

# 按三级标题分割内容
sections = re.split(r'^### ', md_content, flags=re.MULTILINE)[1:]

# 提取每个标题对应的文本内容,去除首尾空白
result = []
for section in sections:
    title_end_pos = section.find('\n')
    content = section[title_end_pos:].strip()
    result.append(content)

# 输出分组结果
for idx, content in enumerate(result, 1):
    print(f"分组{idx}: {repr(content)}")

运行后输出示例:

分组1: 'This is a description'
分组2: 'This is my changelog'
分组3: '- Test 1\n- Test 2'
分组4: ''
分组5: 'Concise summary of what this PR is.'

方法二:命令行 awk 工具提取

适合习惯用命令行的场景,直接处理文件输出分组:

awk '
    /^### / {
        if (group_num > 0) {
            printf "分组%d: %s\n", group_num, content
        }
        group_num++
        content = ""
        next
    }
    {
        if (content != "") content = content "\n" $0
        else content = $0
    }
    END {
        printf "分组%d: %s\n", group_num, content
    }
' your_file.md | sed 's/分组[0-9]*: *$/分组&""//'

执行后会输出和Python脚本一致的结果,空内容的分组会显示为空字符串。

内容的提问来源于stack exchange,提问作者putty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 20:55:20