You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于最长正则匹配将文档分割为目标格式的章节?

按多格式标题分割文档的正则解决方案

需求说明

需要将文档按以下格式的标题分割为独立章节(标题+对应内容):

  • 1. Header(数字序号+标题)
  • I. Header(大写罗马序号+标题)
  • A. Part I. Header(字母序号+Part+罗马序号+标题)
  • Part I. Header(Part+罗马序号+标题)
  • Header(纯单词标题)

示例文本

1. Index
words words
A. Part I. Introduction
words words
4. Information
words words
Conclusion
words words

现有问题

使用的正则表达式包含过多捕获组,且未处理匹配优先级,导致split后返回大量冗余捕获内容,结果不符合预期:

(^(([A-Z]{1}|\d)\.)\s(part (i|ii)(\.\s))?)?(index|introduction|conclusion)$, re.M|re.I

当前错误输出:

['', '1. ', '1.', '1', None, None, None, 'Index', '\nwords words\n', 'A. Part I. ', 'A.', 'A', 'Part I. ', 'I', '. ', 'Introduction', '\nwords words\n', '4. ', '4.', '4', None, None, None, 'Information', '\nwords words\n', None, None, None, None, None, None, 'Conclusion', '\nwords words\n    ']

解决方案

核心思路:用正向预查定位标题起始位置,按最长匹配优先级排列正则规则,避免捕获组干扰分割结果。

实现代码

import re

# 待分割的文档文本
text = """1. Index
words words
A. Part I. Introduction
words words
4. Information
words words
Conclusion
words words"""

# 正则规则:按标题起始位置匹配,最长格式优先
split_pattern = r'(?=^\d+\.\s|^[A-Z]\.\sPart\s[IVX]+\.\s|^Part\s[IVX]+\.\s|^[A-Z]\.\s|^[A-Z][a-z]+$)'

# 执行分割并过滤空内容
sections = re.split(split_pattern, text, flags=re.MULTILINE | re.IGNORECASE)
sections = [sec.strip() for sec in sections if sec.strip()]

print(sections)

正则规则解释

  • (?=...):正向预查,仅匹配标题的起始位置,不消耗字符,避免分割时丢失标题内容
  • ^[A-Z]\.\sPart\s[IVX]+\.\s:优先匹配最长格式的标题(如A. Part I. )
  • ^Part\s[IVX]+\.\s:匹配Part I. 开头的标题
  • ^\d+\.\s:匹配数字序号开头的标题(如1. )
  • ^[A-Z]\.\s:匹配大写字母/罗马序号开头的标题(如I. )
  • ^[A-Z][a-z]+$:匹配纯单词标题(如Conclusion)
  • 规则按长度从长到短排列,确保最长匹配优先,避免短规则提前匹配长标题的部分内容

输出结果

['1. Index\nwords words', 'A. Part I. Introduction\nwords words', '4. Information\nwords words', 'Conclusion\nwords words']

内容的提问来源于stack exchange,提问作者Ainulindalë

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 13:20:29