You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Polars DataFrame转换为Vec<RecordBatch>?

将Polars DataFrame转换为Arrow RecordBatch列表

要把Polars DataFrame转换成Vec<arrow::record_batch::RecordBatch>,核心是用Polars提供的to_arrow_batches方法——这个方法专门为Arrow生态兼容设计,替代你之前尝试的iter_chunks或iter_chunks_physical(这两个返回的是Polars自身的ChunkedArray结构,并非Arrow的RecordBatch)。

步骤说明

  1. 启用Polars的Arrow特性:确保Cargo.toml里的Polars依赖开启arrow特性,转换逻辑需要依赖Arrow相关功能:
[dependencies]
polars = { version = "0.37", features = ["arrow", "lazy"] }
arrow = "47"

(版本号可根据实际情况调整,注意Polars与Arrow的版本兼容性)

  1. 使用to_arrow_batches完成转换:该方法返回Result<Vec<RecordBatch>, PolarsError>,可直接unwrap(生产环境建议做错误处理)。若需自定义批次大小,可通过with_batch_size方法调整。

完整示例代码

use polars::prelude::{DataFrame, df, ToArrowBatches};
use arrow::record_batch::RecordBatch;

fn main() -> Result<(), polars::error::PolarsError> {
    // 创建示例Polars DataFrame
    let polars_df: DataFrame = df!(
        "cat_data"     => &[1.0, 2.0, 3.0, 4.0],
        "dog_data"     => &[1.0, 2.0, 3.0, 4.0],
        "giraffe_data" => &[1.0, 2.0, 3.0, 4.0]
    )?;

    // 转换为Arrow RecordBatch列表,可选自定义批次大小
    let batches: Vec<RecordBatch> = polars_df
        .to_arrow_batches()
        // .with_batch_size(2) // 可选:指定每个批次的行数,默认是Polars的chunk大小
        .collect()?;

    // 遍历输出每个批次
    for batch in batches {
        println!("{:?}", batch);
    }

    Ok(())
}

关键说明

  • to_arrow_batches会自动将Polars列转换为兼容的Arrow数组,再封装成RecordBatch,完美适配Arrow Flight这类依赖Arrow数据结构的场景。
  • 之前尝试的iter_chunks返回的是Polars列的分块迭代器,需要手动将每个ChunkedArray转成Arrow数组再组装成RecordBatch,效率远不如to_arrow_batches直接高效。

内容的提问来源于stack exchange,提问作者user21083829

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 19:06:20