Python拆分大CSV文件及解决TabError缩进错误
拆分大CSV文件的Python实现及缩进错误修复
你的代码报错核心是缩进混用了Tab和空格,Python对缩进要求极严,必须统一用空格(推荐4个)或者Tab,不能混着来,这才触发了TabError。
用Python拆分大CSV的思路没问题,pandas自带的read_csv支持chunksize参数分块读取,不需要额外内置函数,直接用pandas就能搞定,修复并优化后的代码如下:
import pandas as pd # 替换成你的原始CSV文件路径 source_file = "你的大文件路径.csv" # 方案1:按固定行数拆分(比如每6万行一个文件) with pd.read_csv(source_file, chunksize=60000) as reader: for idx, chunk in enumerate(reader): chunk.to_csv(f"拆分文件_{idx}.csv", index=False) # 方案2:精准拆成2个文件 # 先算总行数(减去表头行) total_lines = sum(1 for _ in open(source_file)) - 1 split_line = total_lines // 2 with pd.read_csv(source_file, chunksize=split_line) as reader: for idx, chunk in enumerate(reader): chunk.to_csv(f"拆分文件_{idx}.csv", index=False)
注意事项
- 必须保证缩进统一:
for循环里的chunk.to_csv要和for语句保持相同缩进(全用4个空格),别再混Tab。 index=False:避免写入时自动加索引列,和原文件格式保持一致。- 如果不想用第三方库pandas,也可以用纯Python内置模块实现:
source_file = "你的大文件路径.csv" split_count = 2 # 固定拆成2个文件 with open(source_file, 'r', encoding='utf-8') as f: header = f.readline() lines = f.readlines() line_per_file = len(lines) // split_count for i in range(split_count): start = i * line_per_file # 最后一个文件包含剩余所有行 end = start + line_per_file if i != split_count-1 else None with open(f"拆分文件_{i}.csv", 'w', encoding='utf-8') as out_f: out_f.write(header) out_f.writelines(lines[start:end])
内容的提问来源于stack exchange,提问作者jove
相关产品推荐
相关产品推荐

