You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现大TXT文件分块切割(解决数据丢失问题)

Python 大文本文件分块切割方案(解决数据丢失问题)

问题背景

现有一个每行包含三个值(序号、文件名、固定标识)的大.txt文件,示例内容如下:

1 00000001.setts 0x 
2 00000002.setts 0x 
...
59878 0000e9e6.setts 0x 

文件行数动态(约10万行),需用Python将其切割为每个含1500行的小txt文件,但此前实现时出现数据丢失:仅读取59514行,实际共59878行,求可行方案。

核心问题分析

数据丢失大概率是因为读取文件时未处理好文件末尾的剩余行,或是采用一次性读取全量行的方式引发内存溢出、读取截断问题。

可行实现方案

采用逐行读取、批量写入的方式,既避免一次性加载全量数据占用过多内存,又能确保所有行都被处理。

基础切割代码

def split_large_file(input_file, lines_per_file=1500):
    file_counter = 1
    current_lines = []
    
    # 以只读模式打开原文件,指定编码避免乱码或读取截断
    with open(input_file, 'r', encoding='utf-8') as f:
        # 逐行遍历,确保每一行都被遍历到
        for _, line in enumerate(f):
            # 保留原始行的格式(包括换行符)
            current_lines.append(line)
            
            # 当收集的行数达到指定数量时,写入新文件
            if len(current_lines) == lines_per_file:
                output_filename = f"output_part_{file_counter}.txt"
                with open(output_filename, 'w', encoding='utf-8') as out_f:
                    out_f.writelines(current_lines)
                # 清空当前行列表,准备收集下一批数据
                current_lines = []
                file_counter += 1
        
        # 处理循环结束后剩余的不足1500行的内容
        if current_lines:
            output_filename = f"output_part_{file_counter}.txt"
            with open(output_filename, 'w', encoding='utf-8') as out_f:
                out_f.writelines(current_lines)

# 调用示例,替换为你的大文件路径
split_large_file("your_large_input.txt")

带验证的增强版代码

如果需要确认读取的总行数是否与原文件一致,可以添加计数统计:

def split_large_file(input_file, lines_per_file=1500):
    file_counter = 1
    current_lines = []
    total_read_lines = 0
    
    with open(input_file, 'r', encoding='utf-8') as f:
        for line_num, line in enumerate(f, start=1):
            total_read_lines = line_num
            current_lines.append(line)
            
            if len(current_lines) == lines_per_file:
                output_filename = f"output_part_{file_counter}.txt"
                with open(output_filename, 'w', encoding='utf-8') as out_f:
                    out_f.writelines(current_lines)
                current_lines = []
                file_counter += 1
        
        if current_lines:
            output_filename = f"output_part_{file_counter}.txt"
            with open(output_filename, 'w', encoding='utf-8') as out_f:
                out_f.writelines(current_lines)
    
    # 打印总行数,与原文件实际行数对比验证
    print(f"成功读取总行数: {total_read_lines}")

split_large_file("your_large_input.txt")

关键注意事项

  • 逐行遍历:避免使用readlines()一次性读取所有行,防止大文件导致内存不足或读取截断
  • 处理剩余行:循环结束后必须单独处理未达到批量行数的剩余内容,这是之前数据丢失的常见原因
  • 指定编码:显式设置文件编码(如utf-8),避免因系统默认编码不匹配导致的读取异常
  • 保留原始格式:直接写入读取到的行,不修改内容格式,确保输出文件与原文件一致

内容的提问来源于stack exchange,提问作者schaef

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 13:05:27