You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Polars Rust版本处理大CSV时性能远逊于Python版本?

Polars Rust版CSV读取性能随数据量扩大异常恶化的问题

测试背景与结果

  • 对比Python、Rust生态下的DataFrame库,重点测试跨语言支持的Polars,编写了逻辑完全一致的双语言测试代码
  • 小数据集(2k、1k行的pandas示例CSV):Python与Rust版本性能相近
  • 数据量扩大10倍(20k、10k行):Python版本仅略有变慢,Rust版本耗时却增至原有的8倍

已尝试的优化手段

  • 使用release构建编译Rust代码
  • 减少代码中collect()调用次数
  • 在Cargo.toml中设置lto=fat
  • 替换默认分配器为jemalloc
  • 使用nightly Rust构建
    以上优化均未带来明显性能改善

问题定位

单独测试CSV读取步骤时,发现该环节耗时极高,因此怀疑Polars的CSV读取模块存在bug

测试代码

Python版本

import polars as pl
import time

start = time.time()
df = pl.read_csv("pandas_sample.csv")
# 与Rust版一致的后续数据处理逻辑
end = time.time()
print(f"总耗时: {end - start:.4f}s")

Rust版本

use polars::prelude::*;
use std::time::Instant;

fn main() {
    let start = Instant::now();
    let df = CsvReader::from_path("pandas_sample.csv")
        .unwrap()
        .finish()
        .unwrap();
    // 与Python版一致的后续数据处理逻辑
    let duration = start.elapsed();
    println!("总耗时: {:?}", duration);
}

内容的提问来源于stack exchange,提问作者kelko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 10:35:53