如何使用Python提取文本文件中指定起止标识的特定段落
文本特定段落提取解决方案
原有代码问题梳理
- 遍历列表的同时删除列表元素,会触发索引偏移,导致部分行被跳过无法匹配
- 开头匹配使用严格全等判断,未兼容行尾换行符、额外空格、大小写差异等场景,匹配失败概率高
- 仅完成了开头标记的过滤逻辑,未补充结尾标记
under said Deed of Trust的终止匹配逻辑 - 未实现目标范围内内容的完整保留逻辑
最优实现方案(正则匹配)
直接读取文件全量内容,通过正则一次性匹配起止标记之间的所有内容,代码量少逻辑清晰,非贪婪匹配可满足多重复字段下仅提取目标范围的要求:
import re # 匹配规则:忽略大小写,支持跨行匹配,非贪婪模式避免多结尾干扰 pattern = re.compile(r'SUBSTITUTION OF TRUSTEE.*?under said Deed of Trust', flags=re.IGNORECASE | re.DOTALL) with open('sample.txt', 'rt', encoding='utf-8') as f: content = f.read() match_res = pattern.search(content) if match_res: target_paragraph = match_res.group() print(target_paragraph) else: print("未匹配到目标段落")
针对示例文件运行后,输出结果如下:
SUBSTITUTION OF TRUSTEE AND DEED OF RECONVEYANCE The undersigned, Financial Corporation of Nevada, a Nevada Corporation, as the Owner and Holder of the Note secured by Deed of Trust dated March 1, 2013 made by Elvia Bello, Trustor, to Official Records -- HEREBY substitutes Financial Corporation of Nevada, a Nevada Corporation, as Trustee in lieu of the Trustee therein. Said Note, together with all other indebtedness secured by said Deed of Trust, has been fully paid satisfied; and as successor Trustee, the undersigned does hereby RECONVEY WITHOUT WARRANTY TO THE PERSON OR PERSONS LEGALLY ENTITLED THERETO, all the estate now held by it under said Deed of Trust
大文件适配方案(逐行处理)
如果文件体积过大,一次性读取内存占用过高,可以用标志位逐行过滤,避免内存溢出:
import re start_pattern = re.compile(r'SUBSTITUTION OF TRUSTEE', re.IGNORECASE) end_pattern = re.compile(r'under said Deed of Trust', re.IGNORECASE) target_content = [] in_target = False with open('sample.txt', 'rt', encoding='utf-8') as f: for line in f: # 匹配到开头标记,开启采集 if not in_target and start_pattern.search(line): in_target = True if in_target: target_content.append(line) # 匹配到结尾标记,结束采集 if end_pattern.search(line): break if target_content: print(''.join(target_content)) else: print("未匹配到目标段落")
内容的提问来源于stack exchange,提问作者Sand
相关产品推荐
相关产品推荐

