You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Heading1分割docx/rtf文档?需用Heading1命名且排除其内容

解决方案:分割RTF文件并按Heading1命名(不含Heading1内容)

我刚好折腾过类似的需求,给你一个基于Python的实现方案,完全匹配你的要求——把带Heading1的大RTF拆成多个独立小RTF,每个小文件用对应的Heading1文本当名字,而且文件里不会保留Heading1本身。

步骤1:安装依赖

首先得装处理RTF的专用库,用python-rtf就行,终端里跑这条命令:

pip install python-rtf

步骤2:完整代码实现

from rtf.parser import Parser
from rtf.dumper import Dumper
from rtf.document import Document, Paragraph
import os

def split_rtf_by_heading1(input_path, output_dir):
    # 自动创建输出目录(不存在就新建)
    os.makedirs(output_dir, exist_ok=True)
    
    # 解析输入的大RTF文件
    with open(input_path, 'r', encoding='utf-8') as f:
        parser = Parser(f)
        doc = parser.parse()
    
    current_content = []
    current_filename = None
    
    # 遍历文档里的每一个元素(主要是段落)
    for element in doc.content:
        # 只处理段落类型的内容
        if isinstance(element, Paragraph):
            # 检查这个段落是不是Heading1样式
            is_heading1 = any(style.name == 'Heading1' for style in element.styles)
            
            if is_heading1:
                # 如果之前已经收集了内容,先把上一组内容保存成文件
                if current_content and current_filename:
                    save_rtf(current_content, current_filename, output_dir)
                    current_content = []
                
                # 提取Heading1的文本当文件名,同时替换掉系统不允许的非法字符
                heading_text = ''.join([run.text for run in element.content]).strip()
                # 替换Windows/Unix下的非法文件名字符
                current_filename = heading_text.replace('/', '_').replace('\\', '_').replace(':', '_').replace('*', '_').replace('?', '_').replace('"', '_').replace('<', '_').replace('>', '_').replace('|', '_')
            else:
                # 非Heading1的正文段落,加入当前待保存的内容列表
                current_content.append(element)
    
    # 别忘了保存最后一组内容(最后一个Heading1对应的正文)
    if current_content and current_filename:
        save_rtf(current_content, current_filename, output_dir)

def save_rtf(content, filename, output_dir):
    # 创建一个全新的RTF文档对象
    new_doc = Document()
    # 把收集到的正文段落都加进去
    for para in content:
        new_doc.content.append(para)
    
    # 生成最终的输出文件路径
    output_path = os.path.join(output_dir, f"{filename}.rtf")
    # 写入文件
    with open(output_path, 'w', encoding='utf-8') as f:
        dumper = Dumper(f)
        dumper.dump(new_doc)
    print(f"已生成文件:{output_path}")

# 示例调用(替换成你自己的文件路径和输出文件夹)
if __name__ == "__main__":
    input_rtf = "你的大RTF文件路径.rtf"
    output_folder = "分割后的RTF文件"
    split_rtf_by_heading1(input_rtf, output_folder)

代码说明

  • 样式检测:通过遍历段落的样式属性,精准识别出标记为Heading1的段落
  • 文件名安全处理:自动替换掉所有系统不允许的文件名字符(比如/、:、*这些),避免保存失败
  • 内容精准分割:遇到新的Heading1时,自动收尾上一个文件并开启新的内容收集,确保每个小文件只对应一段Heading1下的正文
  • 纯正文输出:保存的小RTF里完全不包含Heading1文本,只保留对应的正文内容

注意事项

  • 如果你的RTF里的标题样式是中文名称(比如“标题1”),记得把代码里的style.name == 'Heading1'改成对应的样式名
  • 确保原RTF的Heading1是通过样式设置的,而不是手动加粗/调字号的格式,否则可能无法被识别到

内容的提问来源于stack exchange,提问作者JPG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:31:11