You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

用Python合并10000个CSV文件:Pandas过慢求更快方案

快速合并大量单行CSV文件的Python方案

你的Pandas方案慢的核心原因是:每个仅1行的小文件都要触发一次完整的CSV解析流程(类型推断、DataFrame初始化等),10000次重复操作的开销被放大了。下面是几种远快于Pandas的实现方式:

方法一:纯文件IO直接拼接(最快)

利用Python原生文件操作直接读取内容并写入,跳过所有不必要的解析步骤,适合所有文件结构一致的场景:

场景1:所有文件都包含表头(仅保留一份表头)

import glob

output_path = "merged.csv"
csv_files = glob.glob("PATH/*.csv")

with open(output_path, 'w', encoding='utf-8') as outfile:
    # 写入第一个文件的表头+内容
    with open(csv_files[0], 'r', encoding='utf-8') as infile:
        outfile.write(infile.read())
    # 写入剩余文件的内容(跳过表头)
    for file in csv_files[1:]:
        with open(file, 'r', encoding='utf-8') as infile:
            next(infile)  # 跳过表头行
            outfile.write(infile.read())

场景2:所有文件只有数据行(无表头)

import glob

output_path = "merged.csv"
csv_files = glob.glob("PATH/*.csv")

with open(output_path, 'w', encoding='utf-8') as outfile:
    for file in csv_files:
        with open(file, 'r', encoding='utf-8') as infile:
            outfile.write(infile.read())

方法二:使用csv模块(兼顾速度与格式安全性)

如果你的CSV包含特殊字符(如逗号、换行符被引号包裹),纯IO可能会破坏格式,这时用csv模块处理更稳妥,速度依然远快于Pandas:

import glob
import csv

output_path = "merged.csv"
csv_files = glob.glob("PATH/*.csv")

with open(output_path, 'w', encoding='utf-8', newline='') as outfile:
    writer = csv.writer(outfile)
    first_file = True
    for file in csv_files:
        with open(file, 'r', encoding='utf-8', newline='') as infile:
            reader = csv.reader(infile)
            if first_file:
                writer.writerows(reader)
                first_file = False
            else:
                next(reader)  # 跳过表头
                writer.writerows(reader)

方法三:多线程加速IO密集型操作

如果磁盘IO是瓶颈,可以用多线程并行读取文件(注意:多线程适合IO密集场景,CPU密集场景用多进程):

import glob
from concurrent.futures import ThreadPoolExecutor

output_path = "merged.csv"
csv_files = glob.glob("PATH/*.csv")

def read_file(file):
    with open(file, 'r', encoding='utf-8') as f:
        content = f.read()
        # 如果不是第一个文件,去掉表头
        if file != csv_files[0]:
            return content.split('\n', 1)[1] + '\n'
        return content

with ThreadPoolExecutor(max_workers=8) as executor:
    results = executor.map(read_file, csv_files)

with open(output_path, 'w', encoding='utf-8') as outfile:
    outfile.writelines(results)

性能对比

  • 纯IO方法:处理10000个文件通常仅需几秒到十几秒(取决于磁盘速度)
  • csv模块方法:比纯IO慢10%-30%,但格式更安全
  • 原Pandas方法:因重复解析开销,耗时长达数十分钟

内容的提问来源于stack exchange,提问作者physicist1911

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 13:17:18