You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大CSV文件分割咨询:4.000.000行文件拆分方法需求

分割大型CSV文件的实用方法

1. 用Linux/Unix命令行工具(split)

适合快速批量处理,步骤清晰:

  • 先提取CSV表头:
    head -n 1 large_file.csv > header.csv
  • 分割数据部分(跳过表头,每个文件100000行,可按需调整):
    tail -n +2 large_file.csv | split -l 100000 - split_
  • 将表头添加到每个分割后的文件:
    for file in split_*; do cat header.csv "$file" > "$file".csv && rm "$file"; done

2. 用awk命令(简洁高效,自动处理表头)

一行命令完成全部操作,自动为每个小文件追加表头,示例设置每个文件100000行:
awk -v lines=100000 'NR==1{header=$0; next} {file=sprintf("split_%03d.csv", int((NR-2)/lines)+1)} NR==2{print header > file} {print > file}' large_file.csv

3. Python脚本(适合定制化需求)

如果需要按特定字段分割、添加额外处理逻辑,用Python脚本更灵活:

import csv

def split_csv(input_file, output_prefix, lines_per_file=100000):
    with open(input_file, 'r', newline='') as infile:
        reader = csv.reader(infile)
        header = next(reader)
        file_count = 1
        current_writer = None
        current_line = 0
        
        for row in reader:
            if current_line % lines_per_file == 0:
                if current_writer is not None:
                    current_writer.close()
                output_file = f"{output_prefix}_{file_count}.csv"
                outfile = open(output_file, 'w', newline='')
                current_writer = csv.writer(outfile)
                current_writer.writerow(header)
                file_count += 1
            current_writer.writerow(row)
            current_line += 1
        
        if current_writer is not None:
            current_writer.close()

# 调用示例:输入文件名、输出前缀、每个文件行数
split_csv('large_file.csv', 'split', lines_per_file=100000)

内容的提问来源于stack exchange,提问作者Nelson Romero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 02:10:25