You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将1亿行(18.5GB)CSV分割为5万行小Excel?Polars实现遇阻

用Polars分割大型CSV为多份Excel文件的解决方案

问题根源

你的代码里reader.next_batches(1)只取了1个批次,所以循环只执行一次,自然只生成一个Excel文件。另外要注意,next_batches()的参数是批次的数量,不是行数——你之前传50000是错误的,这会试图一次性读取50000个批次,直接导致内存过载或读取失败。

修正后的代码

import os
import polars as pl

def split_csv_to_excel(csv_file_path, output_dir, batch_size=50000):
    # 创建按行分批的CSV读取器,batch_size设置每个文件的行数
    reader = pl.read_csv_batched(csv_file_path, batch_size=batch_size)
    # 自动创建输出目录,已存在也不会报错
    os.makedirs(output_dir, exist_ok=True)
    
    batch_num = 0
    # 循环读取所有批次,直到没有数据
    while True:
        # 每次读取1个批次(对应batch_size行数据)
        current_batches = reader.next_batches(1)
        if not current_batches:
            break
        # 取出批次里的DataFrame
        df = current_batches[0]
        # 生成输出文件路径
        output_path = os.path.join(output_dir, f"batch_{batch_num}.xlsx")
        df.write_excel(output_path)
        batch_num += 1
        print(f"完成文件: {output_path}")

split_csv_to_excel("inputpath", "outputpath")

关键细节

  • os.makedirs(exist_ok=True):简化目录创建逻辑,无需单独判断目录是否存在
  • 循环逻辑:通过while True持续拉取批次,直到next_batches(1)返回空列表,说明所有数据已处理完毕
  • 批次控制:每次只读取1个批次,保证每个Excel文件严格对应batch_size行数据
  • Excel兼容性:.xlsx格式最大支持1048576行,5万行完全在限制内,不用担心溢出

内容的提问来源于stack exchange,提问作者Indratej Reddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 20:22:13