Python提取Fetching分隔文本块并按内容分类写入文件
高效处理大文本文件:按'Fetching'分割块并分类存储
核心思路
因为要处理数十万条条目,绝对不能一次性把整个文件读进内存,必须逐行流式处理——每处理完一个块就立即写入对应文件,内存中只保留当前正在处理的块内容,避免内存溢出。
具体逻辑:
- 初始化空列表存储当前块内容
- 逐行读取原文件:
- 若当前行以'Fetching'开头:
- 如果当前块不为空(说明有未处理的旧块),先判断旧块是否包含
40500000,写入对应文件后清空列表 - 将当前行加入新块
- 如果当前块不为空(说明有未处理的旧块),先判断旧块是否包含
- 否则直接把行加入当前块
- 若当前行以'Fetching'开头:
- 文件读取结束后,处理最后一个剩余的块
完整代码实现
def process_fetching_blocks(input_file_path): current_block = [] # 用with语句自动管理文件句柄,避免资源泄漏 with open('405.txt', 'w', encoding='utf-8') as f405, \ open('Not.txt', 'w', encoding='utf-8') as fnot, \ open(input_file_path, 'r', encoding='utf-8') as fin: for line in fin: # 保留原行的换行符,保证输出格式和原文件一致 if line.strip().startswith('Fetching'): # 处理上一个已完成的块 if current_block: block_content = ''.join(current_block) if '40500000' in block_content: f405.write(block_content) else: fnot.write(block_content) current_block = [] # 开始新块 current_block.append(line) else: current_block.append(line) # 处理文件末尾的最后一个块 if current_block: block_content = ''.join(current_block) if '40500000' in block_content: f405.write(block_content) else: fnot.write(block_content) # 调用示例,替换成你的输入文件路径 process_fetching_blocks('input.txt')
代码说明
- 流式处理:逐行读取输入文件,内存始终只存一个块的内容,哪怕文件几十GB也能稳定处理
- 格式保留:直接保留原行的换行符,输出块和原文件格式完全一致
- 高效写入:批量写入整个块内容,比逐行写入效率更高
- 边界处理:专门处理文件末尾的最后一个块,不会出现遗漏
示例验证
假设输入文件input.txt内容:
Fetching data from source A
Line 1 of block A
40500000 found here
Line 3 of block A
Fetching data from source B
Line 1 of block B
No special code here
Fetching data from source C
40500000 appears again
运行代码后:
405.txt内容:Fetching data from source A Line 1 of block A 40500000 found here Line 3 of block A Fetching data from source C 40500000 appears againNot.txt内容:Fetching data from source B Line 1 of block B No special code here
常见错误思路避坑
很多人会尝试用read()一次性读入整个文件,再用split('Fetching')分割,但这种方式:
- 会把
Fetching本身去掉,需要手动补回,容易出错 - 大文件会直接撑爆内存,完全无法处理数十万条的场景
内容的提问来源于stack exchange,提问作者imp
相关产品推荐
相关产品推荐

