You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Polars中Struct字段从large_string转为dictionary类型?

问题:Polars DataFrame转PyArrow Table匹配历史Schema时的类型兼容问题

原本通过自定义PyArrow Schema将Pandas数据转换为PyArrow格式,用于检测新列或填充空值。现在切换为Polars后,在将DataFrame转为PyArrow Table以匹配历史数据Schema时遇到类型错误。

预定义的PyArrow Schema

import pyarrow as pa

team_info = pa.struct(
    [
        ("_id", pa.string()),
        ("name", pa.string()),
        ("status", pa.dictionary(index_type=pa.int32(), value_type=pa.string())),
    ]
)

schema = pa.schema(
    [
        ("load_timestamp", pa.timestamp(unit="ns", tz="UTC")),
        # 省略其他字段
        ("team_info", team_info),
        # 省略其他字段
    ]
)

遇到的错误

执行return df.to_arrow().cast(schema)转换时,Polars要求team_info嵌套结构中的3个字段全部使用large_string类型,与预定义Schema的类型不匹配。

尝试过的无效方法

曾编写函数试图将嵌套字段status转为Categorical类型,但该操作会新增列,无法原地修改嵌套字段:

def update_nested_status(df: pl.DataFrame, nested_columns: list[str]) -> pl.DataFrame:
    """修复agent_info和monitor_info列的数据类型"""
    cols = [df[col].struct.field("status").cast(pl.Categorical) for col in nested_columns]
    return df.with_columns(cols)

最终可行解决方案

编写了以下函数,能够将Polars数据类型转换为PyArrow Schema中定义的类型,同时按Schema指定的顺序重新排列列:

def align_polars_schema(df: pl.DataFrame, schema: pa.Schema) -> pl.DataFrame:
    """
    将Polars DataFrame的Schema对齐到指定的PyArrow Schema

    参数:
        df: Polars DataFrame
        schema: PyArrow Schema
    """
    schema = pl.from_arrow(schema.empty_table()).schema
    df = df.with_columns([pl.col(col).cast(dtype) for col, dtype in schema.items()])
    return df.select([pl.col(col) for col in schema.keys()])

内容的提问来源于stack exchange,提问作者ldacey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 08:42:56