如何基于Heading1分割docx/rtf文档?需用Heading1命名且排除其内容
解决方案:分割RTF文件并按Heading1命名(不含Heading1内容)
我刚好折腾过类似的需求,给你一个基于Python的实现方案,完全匹配你的要求——把带Heading1的大RTF拆成多个独立小RTF,每个小文件用对应的Heading1文本当名字,而且文件里不会保留Heading1本身。
步骤1:安装依赖
首先得装处理RTF的专用库,用python-rtf就行,终端里跑这条命令:
pip install python-rtf
步骤2:完整代码实现
from rtf.parser import Parser from rtf.dumper import Dumper from rtf.document import Document, Paragraph import os def split_rtf_by_heading1(input_path, output_dir): # 自动创建输出目录(不存在就新建) os.makedirs(output_dir, exist_ok=True) # 解析输入的大RTF文件 with open(input_path, 'r', encoding='utf-8') as f: parser = Parser(f) doc = parser.parse() current_content = [] current_filename = None # 遍历文档里的每一个元素(主要是段落) for element in doc.content: # 只处理段落类型的内容 if isinstance(element, Paragraph): # 检查这个段落是不是Heading1样式 is_heading1 = any(style.name == 'Heading1' for style in element.styles) if is_heading1: # 如果之前已经收集了内容,先把上一组内容保存成文件 if current_content and current_filename: save_rtf(current_content, current_filename, output_dir) current_content = [] # 提取Heading1的文本当文件名,同时替换掉系统不允许的非法字符 heading_text = ''.join([run.text for run in element.content]).strip() # 替换Windows/Unix下的非法文件名字符 current_filename = heading_text.replace('/', '_').replace('\\', '_').replace(':', '_').replace('*', '_').replace('?', '_').replace('"', '_').replace('<', '_').replace('>', '_').replace('|', '_') else: # 非Heading1的正文段落,加入当前待保存的内容列表 current_content.append(element) # 别忘了保存最后一组内容(最后一个Heading1对应的正文) if current_content and current_filename: save_rtf(current_content, current_filename, output_dir) def save_rtf(content, filename, output_dir): # 创建一个全新的RTF文档对象 new_doc = Document() # 把收集到的正文段落都加进去 for para in content: new_doc.content.append(para) # 生成最终的输出文件路径 output_path = os.path.join(output_dir, f"{filename}.rtf") # 写入文件 with open(output_path, 'w', encoding='utf-8') as f: dumper = Dumper(f) dumper.dump(new_doc) print(f"已生成文件:{output_path}") # 示例调用(替换成你自己的文件路径和输出文件夹) if __name__ == "__main__": input_rtf = "你的大RTF文件路径.rtf" output_folder = "分割后的RTF文件" split_rtf_by_heading1(input_rtf, output_folder)
代码说明
- 样式检测:通过遍历段落的样式属性,精准识别出标记为
Heading1的段落 - 文件名安全处理:自动替换掉所有系统不允许的文件名字符(比如
/、:、*这些),避免保存失败 - 内容精准分割:遇到新的Heading1时,自动收尾上一个文件并开启新的内容收集,确保每个小文件只对应一段Heading1下的正文
- 纯正文输出:保存的小RTF里完全不包含Heading1文本,只保留对应的正文内容
注意事项
- 如果你的RTF里的标题样式是中文名称(比如“标题1”),记得把代码里的
style.name == 'Heading1'改成对应的样式名 - 确保原RTF的Heading1是通过样式设置的,而不是手动加粗/调字号的格式,否则可能无法被识别到
内容的提问来源于stack exchange,提问作者JPG
相关产品推荐
相关产品推荐

