如何将Polars中Struct字段从large_string转为dictionary类型?
问题:Polars DataFrame转PyArrow Table匹配历史Schema时的类型兼容问题
原本通过自定义PyArrow Schema将Pandas数据转换为PyArrow格式,用于检测新列或填充空值。现在切换为Polars后,在将DataFrame转为PyArrow Table以匹配历史数据Schema时遇到类型错误。
预定义的PyArrow Schema
import pyarrow as pa team_info = pa.struct( [ ("_id", pa.string()), ("name", pa.string()), ("status", pa.dictionary(index_type=pa.int32(), value_type=pa.string())), ] ) schema = pa.schema( [ ("load_timestamp", pa.timestamp(unit="ns", tz="UTC")), # 省略其他字段 ("team_info", team_info), # 省略其他字段 ] )
遇到的错误
执行return df.to_arrow().cast(schema)转换时,Polars要求team_info嵌套结构中的3个字段全部使用large_string类型,与预定义Schema的类型不匹配。
尝试过的无效方法
曾编写函数试图将嵌套字段status转为Categorical类型,但该操作会新增列,无法原地修改嵌套字段:
def update_nested_status(df: pl.DataFrame, nested_columns: list[str]) -> pl.DataFrame: """修复agent_info和monitor_info列的数据类型""" cols = [df[col].struct.field("status").cast(pl.Categorical) for col in nested_columns] return df.with_columns(cols)
最终可行解决方案
编写了以下函数,能够将Polars数据类型转换为PyArrow Schema中定义的类型,同时按Schema指定的顺序重新排列列:
def align_polars_schema(df: pl.DataFrame, schema: pa.Schema) -> pl.DataFrame: """ 将Polars DataFrame的Schema对齐到指定的PyArrow Schema 参数: df: Polars DataFrame schema: PyArrow Schema """ schema = pl.from_arrow(schema.empty_table()).schema df = df.with_columns([pl.col(col).cast(dtype) for col, dtype in schema.items()]) return df.select([pl.col(col) for col in schema.keys()])
内容的提问来源于stack exchange,提问作者ldacey
相关产品推荐
相关产品推荐

