如何将1亿行(18.5GB)CSV分割为5万行小Excel?Polars实现遇阻
用Polars分割大型CSV为多份Excel文件的解决方案
问题根源
你的代码里reader.next_batches(1)只取了1个批次,所以循环只执行一次,自然只生成一个Excel文件。另外要注意,next_batches()的参数是批次的数量,不是行数——你之前传50000是错误的,这会试图一次性读取50000个批次,直接导致内存过载或读取失败。
修正后的代码
import os import polars as pl def split_csv_to_excel(csv_file_path, output_dir, batch_size=50000): # 创建按行分批的CSV读取器,batch_size设置每个文件的行数 reader = pl.read_csv_batched(csv_file_path, batch_size=batch_size) # 自动创建输出目录,已存在也不会报错 os.makedirs(output_dir, exist_ok=True) batch_num = 0 # 循环读取所有批次,直到没有数据 while True: # 每次读取1个批次(对应batch_size行数据) current_batches = reader.next_batches(1) if not current_batches: break # 取出批次里的DataFrame df = current_batches[0] # 生成输出文件路径 output_path = os.path.join(output_dir, f"batch_{batch_num}.xlsx") df.write_excel(output_path) batch_num += 1 print(f"完成文件: {output_path}") split_csv_to_excel("inputpath", "outputpath")
关键细节
os.makedirs(exist_ok=True):简化目录创建逻辑,无需单独判断目录是否存在- 循环逻辑:通过
while True持续拉取批次,直到next_batches(1)返回空列表,说明所有数据已处理完毕 - 批次控制:每次只读取1个批次,保证每个Excel文件严格对应
batch_size行数据 - Excel兼容性:.xlsx格式最大支持1048576行,5万行完全在限制内,不用担心溢出
内容的提问来源于stack exchange,提问作者Indratej Reddy
相关产品推荐
相关产品推荐

