如何在Python Polars中展开UDF返回值为单独列?
在Polars中展开UDF返回的字典为多列
问题场景
现有如下Python Polars代码,自定义函数do_something返回字典类型数据,调用map_elements后生成包含字典的列,需要将该UDF的返回值展开为单独列,最终得到包含a、b、c三列的DataFrame:
原始代码:
def do_something(text): return {"b": 1, "c": 1} ( pl.DataFrame( { "a": ["x", "y", "z"], } ) .with_columns(pl.col("a").map_elements(do_something)) )
目标输出:
shape: (3, 3) ┌─────┬─────┬─────┐ │ a ┆ b ┆ c │ │ --- ┆ --- ┆ --- │ │ str ┆ i64 ┆ i64 │ ╞═════╪═════╪═════╡ │ x ┆ 1 ┆ 1 │ │ y ┆ 1 ┆ 1 │ │ z ┆ 1 ┆ 1 │ └─────┴─────┴─────┘
解决方案
方法1:指定return_dtype为Struct后展开
通过给map_elements指定return_dtype参数,将字典转换为Polars的Struct类型,再用unnest方法展开为多列:
def do_something(text): return {"b": 1, "c": 1} df = ( pl.DataFrame({"a": ["x", "y", "z"]}) .with_columns( pl.col("a").map_elements( do_something, return_dtype=pl.Struct({"b": pl.Int64, "c": pl.Int64}) ).alias("temp") ) .unnest("temp") ) print(df)
方法2:让UDF直接返回Struct对象
修改自定义函数,直接返回Polars的Struct对象,再通过unnest展开,这种方式性能更优:
def do_something(text): return pl.struct(b=1, c=1) df = ( pl.DataFrame({"a": ["x", "y", "z"]}) .with_columns(pl.col("a").map_elements(do_something).alias("temp")) .unnest("temp") ) print(df)
原理说明
Polars无法直接展开字典类型的列,但Struct类型是Polars原生支持的复合类型,unnest方法可以将Struct的每个字段拆分为单独的列。通过上述两种方式将字典转换为Struct后,即可轻松实现需求。
内容的提问来源于stack exchange,提问作者Laxas
相关产品推荐
相关产品推荐

