You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从多文本文件中匹配PL编号并提取对应区块至独立文件

批量匹配PL编号并提取对应段落的解决方案

核心思路

  1. 先从PL编号文件中提取所有目标PL值,统一格式后存入集合(快速查找)
  2. 按Name分段读取合并文件,每段处理时提取其中的PL值
  3. 若段落PL值匹配目标集合,将整段写入对应独立文件

Python实现代码

# 1. 读取并整理目标PL编号集合
target_pls = set()
with open('pl_numbers.txt', 'r', encoding='utf-8') as pl_file:
    for line in pl_file:
        line = line.strip()
        if not line:
            continue
        # 提取纯PL值,忽略前缀空格、大小写差异
        pl_value = line.split(':', 1)[1].strip()
        target_pls.add(pl_value)

# 2. 分段处理合并文件,提取匹配段落
current_segment = []
with open('merged_file.txt', 'r', encoding='utf-8') as merged_file:
    for line in merged_file:
        stripped_line = line.strip()
        # 遇到新的Name段,先处理上一段内容
        if stripped_line.startswith('Name'):
            if current_segment:
                # 从当前段中提取PL值
                pl_in_segment = None
                for seg_line in current_segment:
                    seg_line_stripped = seg_line.strip()
                    if seg_line_stripped.lower().startswith('pl:'):
                        pl_in_segment = seg_line_stripped.split(':', 1)[1].strip()
                        break
                # 匹配成功则写入文件
                if pl_in_segment and pl_in_segment in target_pls:
                    # 替换小数点避免文件名格式问题
                    filename = f'PL_{pl_in_segment.replace(".", "_")}.txt'
                    with open(filename, 'a', encoding='utf-8') as out_file:
                        out_file.write(''.join(current_segment))
                        out_file.write('\n')
            # 重置当前段,加入新的Name行
            current_segment = [line]
        else:
            current_segment.append(line)
    # 处理文件末尾的最后一段
    if current_segment:
        pl_in_segment = None
        for seg_line in current_segment:
            seg_line_stripped = seg_line.strip()
            if seg_line_stripped.lower().startswith('pl:'):
                pl_in_segment = seg_line_stripped.split(':', 1)[1].strip()
                break
        if pl_in_segment and pl_in_segment in target_pls:
            filename = f'PL_{pl_in_segment.replace(".", "_")}.txt'
            with open(filename, 'a', encoding='utf-8') as out_file:
                out_file.write(''.join(current_segment))
                out_file.write('\n')

关键细节说明

  • 用集合存储目标PL值:4000个元素的查找效率远高于列表,避免循环遍历的性能损耗
  • 格式兼容处理:自动忽略Pl:/pl:大小写、前后空格差异,确保匹配准确性
  • 分段处理逻辑:逐行读取文件,不一次性加载全部内容,支持超大合并文件处理
  • 重复PL处理:同一PL编号对应多个段落时,会追加写入同一个文件,不会覆盖

使用注意事项

  • 将代码与pl_numbers.txt(PL编号文件)、merged_file.txt(合并文本文件)放在同一目录
  • 若文件编码为GBK等非UTF-8格式,需修改代码中encoding参数对应的值

内容的提问来源于stack exchange,提问作者Dsa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 06:22:17