You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python拆分大CSV文件及解决TabError缩进错误

拆分大CSV文件的Python实现及缩进错误修复

你的代码报错核心是缩进混用了Tab和空格,Python对缩进要求极严,必须统一用空格(推荐4个)或者Tab,不能混着来,这才触发了TabError。

用Python拆分大CSV的思路没问题,pandas自带的read_csv支持chunksize参数分块读取,不需要额外内置函数,直接用pandas就能搞定,修复并优化后的代码如下:

import pandas as pd

# 替换成你的原始CSV文件路径
source_file = "你的大文件路径.csv"

# 方案1:按固定行数拆分(比如每6万行一个文件)
with pd.read_csv(source_file, chunksize=60000) as reader:
    for idx, chunk in enumerate(reader):
        chunk.to_csv(f"拆分文件_{idx}.csv", index=False)

# 方案2:精准拆成2个文件
# 先算总行数(减去表头行)
total_lines = sum(1 for _ in open(source_file)) - 1
split_line = total_lines // 2

with pd.read_csv(source_file, chunksize=split_line) as reader:
    for idx, chunk in enumerate(reader):
        chunk.to_csv(f"拆分文件_{idx}.csv", index=False)

注意事项

  • 必须保证缩进统一:for循环里的chunk.to_csv要和for语句保持相同缩进(全用4个空格),别再混Tab。
  • index=False:避免写入时自动加索引列,和原文件格式保持一致。
  • 如果不想用第三方库pandas,也可以用纯Python内置模块实现:
source_file = "你的大文件路径.csv"
split_count = 2  # 固定拆成2个文件

with open(source_file, 'r', encoding='utf-8') as f:
    header = f.readline()
    lines = f.readlines()
    line_per_file = len(lines) // split_count

    for i in range(split_count):
        start = i * line_per_file
        # 最后一个文件包含剩余所有行
        end = start + line_per_file if i != split_count-1 else None
        with open(f"拆分文件_{i}.csv", 'w', encoding='utf-8') as out_f:
            out_f.write(header)
            out_f.writelines(lines[start:end])

内容的提问来源于stack exchange,提问作者jove

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 03:47:48