Python实现大TXT文件分块切割(解决数据丢失问题)
Python 大文本文件分块切割方案(解决数据丢失问题)
问题背景
现有一个每行包含三个值(序号、文件名、固定标识)的大.txt文件,示例内容如下:
1 00000001.setts 0x 2 00000002.setts 0x ... 59878 0000e9e6.setts 0x文件行数动态(约10万行),需用Python将其切割为每个含1500行的小txt文件,但此前实现时出现数据丢失:仅读取59514行,实际共59878行,求可行方案。
核心问题分析
数据丢失大概率是因为读取文件时未处理好文件末尾的剩余行,或是采用一次性读取全量行的方式引发内存溢出、读取截断问题。
可行实现方案
采用逐行读取、批量写入的方式,既避免一次性加载全量数据占用过多内存,又能确保所有行都被处理。
基础切割代码
def split_large_file(input_file, lines_per_file=1500): file_counter = 1 current_lines = [] # 以只读模式打开原文件,指定编码避免乱码或读取截断 with open(input_file, 'r', encoding='utf-8') as f: # 逐行遍历,确保每一行都被遍历到 for _, line in enumerate(f): # 保留原始行的格式(包括换行符) current_lines.append(line) # 当收集的行数达到指定数量时,写入新文件 if len(current_lines) == lines_per_file: output_filename = f"output_part_{file_counter}.txt" with open(output_filename, 'w', encoding='utf-8') as out_f: out_f.writelines(current_lines) # 清空当前行列表,准备收集下一批数据 current_lines = [] file_counter += 1 # 处理循环结束后剩余的不足1500行的内容 if current_lines: output_filename = f"output_part_{file_counter}.txt" with open(output_filename, 'w', encoding='utf-8') as out_f: out_f.writelines(current_lines) # 调用示例,替换为你的大文件路径 split_large_file("your_large_input.txt")
带验证的增强版代码
如果需要确认读取的总行数是否与原文件一致,可以添加计数统计:
def split_large_file(input_file, lines_per_file=1500): file_counter = 1 current_lines = [] total_read_lines = 0 with open(input_file, 'r', encoding='utf-8') as f: for line_num, line in enumerate(f, start=1): total_read_lines = line_num current_lines.append(line) if len(current_lines) == lines_per_file: output_filename = f"output_part_{file_counter}.txt" with open(output_filename, 'w', encoding='utf-8') as out_f: out_f.writelines(current_lines) current_lines = [] file_counter += 1 if current_lines: output_filename = f"output_part_{file_counter}.txt" with open(output_filename, 'w', encoding='utf-8') as out_f: out_f.writelines(current_lines) # 打印总行数,与原文件实际行数对比验证 print(f"成功读取总行数: {total_read_lines}") split_large_file("your_large_input.txt")
关键注意事项
- 逐行遍历:避免使用
readlines()一次性读取所有行,防止大文件导致内存不足或读取截断 - 处理剩余行:循环结束后必须单独处理未达到批量行数的剩余内容,这是之前数据丢失的常见原因
- 指定编码:显式设置文件编码(如
utf-8),避免因系统默认编码不匹配导致的读取异常 - 保留原始格式:直接写入读取到的行,不修改内容格式,确保输出文件与原文件一致
内容的提问来源于stack exchange,提问作者schaef
相关产品推荐
相关产品推荐

