You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Polars保存带分区Parquet文件时分区名异常问题求助

解决Polars保存分区Parquet时分区名带方括号的问题

当你用Polars结合use_pyarrow=True保存分区Parquet文件时,会遇到分区目录名被加上方括号(比如[animals=cat])的问题,导致命名不符合常规的Hive风格(应该是animals=cat)。

问题原因

这是因为PyArrow默认的分区命名格式不是Hive风格,Polars通过pyarrow_options传递参数时,需要显式指定分区的命名规则。

解决方案

在pyarrow_options中添加partitioning参数,指定为Hive风格即可解决。

修改后的代码示例

output_path = './data_parquet'
data = {'animals':['cat','dog','cat','dog','cat','dog','cat','dog']}
pl.DataFrame(data).write_parquet(
    output_path,
    use_pyarrow=True,
    pyarrow_options={
        "partition_cols": ['animals'],
        "partitioning": "hive"  # 指定Hive风格分区命名
    }
)

另一种更明确的写法

也可以直接使用PyArrow的HivePartitioning类来指定:

import pyarrow as pa

output_path = './data_parquet'
data = {'animals':['cat','dog','cat','dog','cat','dog','cat','dog']}
pl.DataFrame(data).write_parquet(
    output_path,
    use_pyarrow=True,
    pyarrow_options={
        "partition_cols": ['animals'],
        "partitioning": pa.partitioning.HivePartitioning()
    }
)

添加以上参数后,PyArrow会生成标准的Hive格式分区目录(如animals=cat),不会再出现多余的方括号。

内容的提问来源于stack exchange,提问作者Denis Lvov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 04:23:15