如何从多文本文件中匹配PL编号并提取对应区块至独立文件
批量匹配PL编号并提取对应段落的解决方案
核心思路
- 先从PL编号文件中提取所有目标PL值,统一格式后存入集合(快速查找)
- 按
Name分段读取合并文件,每段处理时提取其中的PL值 - 若段落PL值匹配目标集合,将整段写入对应独立文件
Python实现代码
# 1. 读取并整理目标PL编号集合 target_pls = set() with open('pl_numbers.txt', 'r', encoding='utf-8') as pl_file: for line in pl_file: line = line.strip() if not line: continue # 提取纯PL值,忽略前缀空格、大小写差异 pl_value = line.split(':', 1)[1].strip() target_pls.add(pl_value) # 2. 分段处理合并文件,提取匹配段落 current_segment = [] with open('merged_file.txt', 'r', encoding='utf-8') as merged_file: for line in merged_file: stripped_line = line.strip() # 遇到新的Name段,先处理上一段内容 if stripped_line.startswith('Name'): if current_segment: # 从当前段中提取PL值 pl_in_segment = None for seg_line in current_segment: seg_line_stripped = seg_line.strip() if seg_line_stripped.lower().startswith('pl:'): pl_in_segment = seg_line_stripped.split(':', 1)[1].strip() break # 匹配成功则写入文件 if pl_in_segment and pl_in_segment in target_pls: # 替换小数点避免文件名格式问题 filename = f'PL_{pl_in_segment.replace(".", "_")}.txt' with open(filename, 'a', encoding='utf-8') as out_file: out_file.write(''.join(current_segment)) out_file.write('\n') # 重置当前段,加入新的Name行 current_segment = [line] else: current_segment.append(line) # 处理文件末尾的最后一段 if current_segment: pl_in_segment = None for seg_line in current_segment: seg_line_stripped = seg_line.strip() if seg_line_stripped.lower().startswith('pl:'): pl_in_segment = seg_line_stripped.split(':', 1)[1].strip() break if pl_in_segment and pl_in_segment in target_pls: filename = f'PL_{pl_in_segment.replace(".", "_")}.txt' with open(filename, 'a', encoding='utf-8') as out_file: out_file.write(''.join(current_segment)) out_file.write('\n')
关键细节说明
- 用集合存储目标PL值:4000个元素的查找效率远高于列表,避免循环遍历的性能损耗
- 格式兼容处理:自动忽略
Pl:/pl:大小写、前后空格差异,确保匹配准确性 - 分段处理逻辑:逐行读取文件,不一次性加载全部内容,支持超大合并文件处理
- 重复PL处理:同一PL编号对应多个段落时,会追加写入同一个文件,不会覆盖
使用注意事项
- 将代码与
pl_numbers.txt(PL编号文件)、merged_file.txt(合并文本文件)放在同一目录 - 若文件编码为GBK等非UTF-8格式,需修改代码中
encoding参数对应的值
内容的提问来源于stack exchange,提问作者Dsa
相关产品推荐
相关产品推荐

