如何用Python高效排序大CSV文件(按日期,不加载全量进内存)
没问题!这种超大CSV文件的排序需求,核心思路就是外部排序——把大文件拆成内存能hold住的小片段,分别排序后再合并成最终的有序文件。结合你的场景,我给你一套纯Python实现的高效方案,完全不用加载整个文件,还能精准按旧日期在前排序:
核心思路拆解
- 拆分并排序小文件:把8GB的大文件切成多个内存可容纳的小块(比如每次读100万行),对每个小块按日期字段排序后存为临时文件。为了提升排序效率,我们会把
mm/dd/yyyy格式的日期转成yyyy-mm-dd的字符串(这样直接按字符串排序就等价于日期排序,不用反复解析datetime对象)。 - 多路归并临时文件:用堆结构(Python的
heapq模块)把所有有序的临时文件合并成一个大的有序文件,每次只从每个临时文件读取一行,完全不会占用过多内存。
完整代码实现
第一步:拆分大文件为有序小文件
import os import tempfile from datetime import datetime def split_and_sort_large_file(input_file, chunk_size=1_000_000): temp_files = [] # 用更大的缓冲区减少IO次数,提升读取速度 with open(input_file, 'r', encoding='utf-8', buffering=1024*1024*10) as f: chunk = [] for line in f: line = line.strip() if not line: continue # 拆分字段:第一字段、日期、剩余内容(假设第一字段无逗号,按用户描述处理) parts = line.split(',', 2) if len(parts) < 2: # 跳过无效行(比如字段不足的情况) continue date_str = parts[1] try: # 把mm/dd/yyyy转成yyyy-mm-dd,方便字符串排序 dt = datetime.strptime(date_str, '%m/%d/%Y') sortable_date = dt.strftime('%Y-%m-%d') # 把可排序的日期放在行首,后面接原行内容 modified_line = f"{sortable_date},{line}" chunk.append(modified_line) except ValueError: # 日期格式错误的行,可根据需求改为写入错误日志 continue # 当块大小达到设定值时,排序并保存为临时文件 if len(chunk) >= chunk_size: chunk.sort() # 创建临时文件,自动生成唯一路径 temp_fd, temp_path = tempfile.mkstemp(suffix='.csv') with os.fdopen(temp_fd, 'w', encoding='utf-8') as temp_f: temp_f.write('\n'.join(chunk) + '\n') temp_files.append(temp_path) chunk = [] # 处理最后一块剩余数据 if chunk: chunk.sort() temp_fd, temp_path = tempfile.mkstemp(suffix='.csv') with os.fdopen(temp_fd, 'w', encoding='utf-8') as temp_f: temp_f.write('\n'.join(chunk) + '\n') temp_files.append(temp_path) return temp_files
第二步:归并所有有序临时文件
import heapq def merge_sorted_files(temp_files, output_file): # 为每个临时文件创建迭代器,返回(排序键, 原行内容) file_iterators = [] for temp_path in temp_files: f = open(temp_path, 'r', encoding='utf-8', buffering=1024*1024*10) # 生成器函数,逐行读取并拆分排序键和原行 def line_generator(fp): for line in fp: line = line.strip() if not line: continue sort_key, original_line = line.split(',', 1) yield (sort_key, original_line) file_iterators.append(line_generator(f)) # 用heapq.merge合并多个有序迭代器,自动按sort_key排序(旧日期在前) with open(output_file, 'w', encoding='utf-8', buffering=1024*1024*10) as out_f: for sort_key, original_line in heapq.merge(*file_iterators): out_f.write(original_line + '\n') # 清理临时文件:关闭文件句柄并删除文件 # 从生成器闭包中取出文件对象关闭 for gen in file_iterators: gen.__closure__[0].cell_contents.close() for temp_path in temp_files: os.remove(temp_path)
主函数调用
def main(): input_csv = 'your_large_file.csv' # 替换成你的输入文件路径 output_csv = 'sorted_result.csv' # 替换成你的输出文件路径 # 根据内存调整chunk_size:4GB内存建议设为100万-200万行,可测试后优化 temp_files = split_and_sort_large_file(input_csv, chunk_size=1_000_000) merge_sorted_files(temp_files, output_csv) print(f"排序完成!结果已保存到 {output_csv}") if __name__ == '__main__': main()
关键优化点说明
- 日期预处理:把
mm/dd/yyyy转成yyyy-mm-dd字符串,直接用字符串排序替代datetime解析,大幅提升排序速度。 - 缓冲区优化:设置
buffering参数为10MB,减少磁盘IO的次数,提升文件读写效率。 - 内存控制:通过
chunk_size控制每次加载到内存的行数,4GB内存下100万行是比较合理的起点,可根据实际测试调整(比如内存剩余多就调大,反之调小)。 - 自动清理临时文件:归并完成后自动删除临时文件,避免占用磁盘空间。
注意事项
- 如果你的第一字段包含逗号(虽然你描述是无引号文本字段),那
split(',',2)会出错,这时候需要用csv.reader来读取每行,确保正确拆分字段。 - 日期格式错误的行当前是直接跳过,你可以修改代码把这些行写入单独的错误日志文件,方便后续排查。
内容的提问来源于stack exchange,提问作者Forinstance
相关产品推荐
相关产品推荐

