用Python合并10000个CSV文件:Pandas过慢求更快方案
快速合并大量单行CSV文件的Python方案
你的Pandas方案慢的核心原因是:每个仅1行的小文件都要触发一次完整的CSV解析流程(类型推断、DataFrame初始化等),10000次重复操作的开销被放大了。下面是几种远快于Pandas的实现方式:
方法一:纯文件IO直接拼接(最快)
利用Python原生文件操作直接读取内容并写入,跳过所有不必要的解析步骤,适合所有文件结构一致的场景:
场景1:所有文件都包含表头(仅保留一份表头)
import glob output_path = "merged.csv" csv_files = glob.glob("PATH/*.csv") with open(output_path, 'w', encoding='utf-8') as outfile: # 写入第一个文件的表头+内容 with open(csv_files[0], 'r', encoding='utf-8') as infile: outfile.write(infile.read()) # 写入剩余文件的内容(跳过表头) for file in csv_files[1:]: with open(file, 'r', encoding='utf-8') as infile: next(infile) # 跳过表头行 outfile.write(infile.read())
场景2:所有文件只有数据行(无表头)
import glob output_path = "merged.csv" csv_files = glob.glob("PATH/*.csv") with open(output_path, 'w', encoding='utf-8') as outfile: for file in csv_files: with open(file, 'r', encoding='utf-8') as infile: outfile.write(infile.read())
方法二:使用csv模块(兼顾速度与格式安全性)
如果你的CSV包含特殊字符(如逗号、换行符被引号包裹),纯IO可能会破坏格式,这时用csv模块处理更稳妥,速度依然远快于Pandas:
import glob import csv output_path = "merged.csv" csv_files = glob.glob("PATH/*.csv") with open(output_path, 'w', encoding='utf-8', newline='') as outfile: writer = csv.writer(outfile) first_file = True for file in csv_files: with open(file, 'r', encoding='utf-8', newline='') as infile: reader = csv.reader(infile) if first_file: writer.writerows(reader) first_file = False else: next(reader) # 跳过表头 writer.writerows(reader)
方法三:多线程加速IO密集型操作
如果磁盘IO是瓶颈,可以用多线程并行读取文件(注意:多线程适合IO密集场景,CPU密集场景用多进程):
import glob from concurrent.futures import ThreadPoolExecutor output_path = "merged.csv" csv_files = glob.glob("PATH/*.csv") def read_file(file): with open(file, 'r', encoding='utf-8') as f: content = f.read() # 如果不是第一个文件,去掉表头 if file != csv_files[0]: return content.split('\n', 1)[1] + '\n' return content with ThreadPoolExecutor(max_workers=8) as executor: results = executor.map(read_file, csv_files) with open(output_path, 'w', encoding='utf-8') as outfile: outfile.writelines(results)
性能对比
- 纯IO方法:处理10000个文件通常仅需几秒到十几秒(取决于磁盘速度)
- csv模块方法:比纯IO慢10%-30%,但格式更安全
- 原Pandas方法:因重复解析开销,耗时长达数十分钟
内容的提问来源于stack exchange,提问作者physicist1911
相关产品推荐
相关产品推荐

