You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取Fetching分隔文本块并按内容分类写入文件

高效处理大文本文件:按'Fetching'分割块并分类存储

核心思路

因为要处理数十万条条目,绝对不能一次性把整个文件读进内存,必须逐行流式处理——每处理完一个块就立即写入对应文件,内存中只保留当前正在处理的块内容,避免内存溢出。

具体逻辑:

  • 初始化空列表存储当前块内容
  • 逐行读取原文件:
    • 若当前行以'Fetching'开头:
      • 如果当前块不为空(说明有未处理的旧块),先判断旧块是否包含40500000,写入对应文件后清空列表
      • 将当前行加入新块
    • 否则直接把行加入当前块
  • 文件读取结束后,处理最后一个剩余的块

完整代码实现

def process_fetching_blocks(input_file_path):
    current_block = []
    # 用with语句自动管理文件句柄,避免资源泄漏
    with open('405.txt', 'w', encoding='utf-8') as f405, \
         open('Not.txt', 'w', encoding='utf-8') as fnot, \
         open(input_file_path, 'r', encoding='utf-8') as fin:
        
        for line in fin:
            # 保留原行的换行符,保证输出格式和原文件一致
            if line.strip().startswith('Fetching'):
                # 处理上一个已完成的块
                if current_block:
                    block_content = ''.join(current_block)
                    if '40500000' in block_content:
                        f405.write(block_content)
                    else:
                        fnot.write(block_content)
                    current_block = []
                # 开始新块
                current_block.append(line)
            else:
                current_block.append(line)
        
        # 处理文件末尾的最后一个块
        if current_block:
            block_content = ''.join(current_block)
            if '40500000' in block_content:
                f405.write(block_content)
            else:
                fnot.write(block_content)

# 调用示例,替换成你的输入文件路径
process_fetching_blocks('input.txt')

代码说明

  1. 流式处理:逐行读取输入文件,内存始终只存一个块的内容,哪怕文件几十GB也能稳定处理
  2. 格式保留:直接保留原行的换行符,输出块和原文件格式完全一致
  3. 高效写入:批量写入整个块内容,比逐行写入效率更高
  4. 边界处理:专门处理文件末尾的最后一个块,不会出现遗漏

示例验证

假设输入文件input.txt内容:

Fetching data from source A
Line 1 of block A
40500000 found here
Line 3 of block A
Fetching data from source B
Line 1 of block B
No special code here
Fetching data from source C
40500000 appears again

运行代码后:

  • 405.txt内容:
    Fetching data from source A
    Line 1 of block A
    40500000 found here
    Line 3 of block A
    Fetching data from source C
    40500000 appears again
    
  • Not.txt内容:
    Fetching data from source B
    Line 1 of block B
    No special code here
    

常见错误思路避坑

很多人会尝试用read()一次性读入整个文件,再用split('Fetching')分割,但这种方式:

  • 会把Fetching本身去掉,需要手动补回,容易出错
  • 大文件会直接撑爆内存,完全无法处理数十万条的场景

内容的提问来源于stack exchange,提问作者imp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 02:58:39