如何用正则表达式提取带点分节号(1.1、1.1.1等)的章节间文本?
解决方案
核心思路
你的问题出在使用了贪婪匹配(.*),它会直接匹配到文档最后一个符合条件的章节号。要解决这个问题,需要用非贪婪匹配结合通用章节号正则模式,精准匹配目标章节到下一个章节(或文档结尾)之间的内容。
通用正则表达式
要提取指定章节(比如1.1)的内容,使用以下正则:
(?s)^1\.1.*?(?=^\d+(\.\d+)*|\Z)
正则各部分解释:
(?s):开启单行模式,让.可以匹配换行符,适配跨多行的章节内容^1\.1:匹配行首的目标章节号(替换成你需要提取的分节号,注意转义圆点).*?:非贪婪匹配任意内容,确保匹配到第一个后续章节就停止(?=^\d+(\.\d+)*|\Z):正向预查,匹配两种终止条件:^\d+(\.\d+)*:行首的任意层级章节号(比如1.1.1、5.1等)\Z:文档结尾,处理最后一个章节的内容提取
示例用法(以Python为例)
假设你的文档内容存在变量doc_text中,提取1.1章节的内容:
import re doc_text = """1.1 Introduction There are some sentences in here that I want and I want to do other things with them. There could be hundreds of sentences, who cares. 1.1.1 Something Else This is where we talk about something else in life. ... 5.1.1 Conclusion """ target_section = "1.1" # 转义分节号里的圆点,避免正则语法冲突 escaped_section = re.escape(target_section) pattern = rf"(?s)^{escaped_section}.*?(?=^\d+(\.\d+)*|\Z)" match = re.search(pattern, doc_text, re.MULTILINE) if match: content = match.group() # 可以进一步去掉章节标题,只保留正文 content_without_title = re.sub(rf"^{escaped_section}.*?\n", "", content, flags=re.MULTILINE) print(content_without_title)
批量处理目录分节号
如果已经拿到目录的分节号列表(比如["1.1", "1.1.1", "5.1.1"]),可以循环处理每个分节号,匹配它和下一个分节号之间的内容:
sections = ["1.1", "1.1.1", "5.1.1"] for i in range(len(sections)): current_section = sections[i] if i < len(sections)-1: next_section = sections[i+1] escaped_current = re.escape(current_section) escaped_next = re.escape(next_section) pattern = rf"(?s)^{escaped_current}.*?(?=^{escaped_next})" else: # 最后一个章节,匹配到文档结尾 escaped_current = re.escape(current_section) pattern = rf"(?s)^{escaped_current}.*?(?=\Z)" match = re.search(pattern, doc_text, re.MULTILINE) if match: # 处理内容... print(f"=== 章节 {current_section} 内容 ===") print(match.group())
内容的提问来源于stack exchange,提问作者kpcrash
相关产品推荐
相关产品推荐

