You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python按行值拆分NEM12格式大型CSV文件为多个小文件问题咨询

背景

我有一个符合特定格式(NEM12)的大型CSV文件,因体积过大无法正常处理,该文件格式规则如下:

  • 文件始终以100开头
  • 首字段为200的行代表新数据集的起始
  • 首字段为300或400的行属于对应数据集的内容
  • 文件始终以900结尾

格式示例:

100 NEM12               
200 NMI INFO    INFO        
300 20211001    0   0   0   0
400 20  20  F17     
300 20211002    0   0   0   0
300 20211003    0   0   0   0
200 NMI INFO    INFO        
300 20211001    0   0   0   0
300 20211002    0   0   0   0
300 20211003    0   0   0   0
300 20211004    0   0   0   0
300 20211005    0   0   0   0
…                   
200 NMI INFO    INFO        
300 20211001    0   0   0   0
300 20211002    0   0   0   0
400 20  20  F17     
300 20211003    0   0   0   0
300 20211004    0   0   0   0
900 
需求

将该大型文件拆分为数百个小文件,每个小文件仅包含一条200行及其对应的所有300、400行数据,且保留NEM12格式要求的100开头、900结尾规则。

原有方案问题

尝试用pandas读取文件因行列结构不规则失败;自行编写的逐行遍历代码存在三个核心问题:

  1. Python原生无left方法,直接调用会报错
  2. csv.writer的writerow方法接收可迭代对象,直接传入字符串会将每个字符拆分为单独列,比如200会被拆为['2','0','0']三个字段写入
  3. 频繁打开关闭文件会大幅降低大文件处理效率
解决方案

直接用原生Python逐行读写即可,不需要引入额外依赖,内存占用极低,适合处理超大NEM12文件:

import os
# 配置参数
input_file_path = "你的大文件路径.csv"
output_dir = "./split_nem12_files" # 拆分后文件存放目录
# 创建输出目录
os.makedirs(output_dir, exist_ok=True)
current_file = None
first_line = None
# 逐行读取大文件
with open(input_file_path, 'r', encoding='utf-8') as f:
    for line in f:
        line_stripped = line.strip()
        if not line_stripped:
            continue
        # 读取首行100作为每个小文件的开头
        if line_stripped.startswith('100'):
            first_line = line
            continue
        # 遇到900结尾行,关闭当前打开的小文件
        if line_stripped.startswith('900'):
            if current_file:
                current_file.write('900\n')
                current_file.close()
            continue
        # 遇到200开头行,创建新的小文件
        if line_stripped.startswith('200'):
            # 先关闭上一个打开的小文件,写入900结尾
            if current_file:
                current_file.write('900\n')
                current_file.close()
            # 生成新文件名,可根据需要调整,这里取200行内容替换特殊字符作为文件名
            fname = line_stripped.replace(',', '_').replace(' ', '_').replace('/', '_') + '.csv'
            full_path = os.path.join(output_dir, fname)
            current_file = open(full_path, 'w', encoding='utf-8')
            # 写入100开头和当前200行
            current_file.write(first_line)
            current_file.write(line)
            continue
        # 遇到300/400行,写入当前打开的小文件
        if line_stripped.startswith(('300', '400')):
            if current_file:
                current_file.write(line)
# 收尾,确保最后一个文件被正常关闭
if current_file and not current_file.closed:
    current_file.write('900\n')
    current_file.close()

方案说明

  1. 全程逐行读写,内存仅保留当前行和第一行100内容,不管原文件多大都能正常运行
  2. 直接写入原行内容,避免用csv.writer的列拆分问题,完全保留原文件格式
  3. 每个小文件自动补全100开头和900结尾,符合NEM12标准格式
  4. 自动创建输出目录,不会因为目录不存在报错

内容的提问来源于stack exchange,提问作者Bobby Heyer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 20:24:01