如何使用Python按指定行索引数组拆分CSV文件
大CSV文件按指定行索引拆分实现方案
原代码存在的问题
- 直接调用
readlines()将300万行全量内容加载到内存,资源占用高,大文件场景易触发内存溢出 - 循环遍历对象错误:遍历了整个CSV的行长度,而非拆分索引列表,会生成远多于预期的文件
- 切片逻辑错误:
csvfile[index[i]:1+index[i]]仅能取到单一行数据,不符合按区间拆分的需求
实现方案
方案1:内存友好逐行读取版(推荐300万行场景使用)
无需加载全量文件,内存占用稳定在100M以内,运行无内存压力:
# 替换为你自己的拆分索引列表 split_index = [0, 1000, 5000, 20000] # 预生成拆分区间,+1是为了匹配原文件跳过表头后的行号 split_intervals = [(split_index[i]+1, split_index[i+1]+1) for i in range(len(split_index)-1)] # 单独读取表头 with open('input_0.csv', 'r', encoding='utf-8') as f: header = f.readline() file_count = 1 current_start, current_end = split_intervals[0] # 初始化第一个输出文件 out_file = open(f"{file_count}.csv", "w", encoding="utf-8") out_file.write(header) with open('input_0.csv', 'r', encoding='utf-8') as f: # 跳过表头行 next(f) for line_num, line in enumerate(f): # 超出当前区间则切换输出文件 if line_num > current_end: out_file.close() file_count += 1 if file_count - 1 >= len(split_intervals): break current_start, current_end = split_intervals[file_count - 1] out_file = open(f"{file_count}.csv", "w", encoding="utf-8") out_file.write(header) # 符合区间则写入数据 if current_start <= line_num <= current_end: out_file.write(line) # 关闭最后一个输出文件 out_file.close()
方案2:全量加载简化版(内存足够时使用)
代码逻辑更简洁,适合内存充足的场景:
# 替换为你自己的拆分索引列表 split_index = [0, 1000, 5000, 20000] with open('input_0.csv', 'r', encoding='utf-8') as f: all_lines = f.readlines() header = all_lines[0] for idx in range(len(split_index) - 1): # 计算对应原文件的行切片范围 start = split_index[idx] + 1 end = split_index[idx + 1] + 1 output_content = [header] + all_lines[start:end + 1] with open(f"{idx + 1}.csv", "w", encoding="utf-8") as f: f.writelines(output_content)
注意事项
- 若你给出的拆分索引已包含表头行号,删除代码中所有对应位置的
+1逻辑即可 - 若CSV文件为GBK编码,将
encoding='utf-8'修改为encoding='gbk'即可
内容的提问来源于stack exchange,提问作者MMM
相关产品推荐
相关产品推荐

