如何用返回元组的函数为Polars DataFrame添加多列并指定类型
问题场景
现有单列Polars DataFrame:
df = pl.DataFrame({ 'a':[1,2,3,4] })
结构如下:
shape: (4, 1) ┌─────┐ │ a │ │ --- │ │ i64 │ ╞═════╡ │ 1 │ │ 2 │ │ 3 │ │ 4 │ └─────┘
另有函数f接收整数参数,返回(int, int)元组:
def f(i): return (i*10, i*100)
需调用该函数为DataFrame添加b、c两列,且显式指定列的数据类型(实际场景可能为Polars结构体、数组等复杂类型)。
尝试以下代码未达预期:
df.with_columns(pl.col('a').map_elements(f).alias('temp'))
得到结果:
┌─────┬───────────┐ │ a ┆ temp │ │ --- ┆ --- │ │ i64 ┆ object │ ╞═════╪═══════════╡ │ 1 ┆ (10, 100) │ │ 2 ┆ (20, 200) │ │ 3 ┆ (30, 300) │ │ 4 ┆ (40, 400) │ └─────┴───────────┘
期望结果:
shape: (4, 3) ┌─────┬─────┬─────┐ │ a ┆ b ┆ c │ │ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ i64 │ ╞═════╪═════╪═════╡ │ 1 ┆ 10 ┆ 100 │ │ 2 ┆ 20 ┆ 200 │ │ 3 ┆ 30 ┆ 300 │ │ 4 ┆ 40 ┆ 400 │ └─────┴─────┴─────┘
解决方案
方法1:通过Struct类型显式定义并展开
利用map_elements时指定返回类型为Polars Struct,再将Struct拆分为独立列,精准控制每列数据类型:
import polars as pl df = pl.DataFrame({'a': [1,2,3,4]}) def f(i): return (i*10, i*100) result_df = df.with_columns( pl.col('a').map_elements( f, return_dtype=pl.Struct([pl.Field('b', pl.Int64), pl.Field('c', pl.Int64)]) ).alias('temp') ).unnest('temp') print(result_df)
执行后结果与预期一致,b、c列类型被显式指定为Int64。
方法2:使用map_batches批量处理(高性能)
数据量较大时,推荐用map_batches替代map_elements,性能更优,同时支持显式指定返回类型:
import polars as pl df = pl.DataFrame({'a': [1,2,3,4]}) def f(i): return (i*10, i*100) def batch_f(series: pl.Series) -> pl.DataFrame: tuples = series.map_elements(f) return pl.DataFrame(tuples.to_list(), schema={'b': pl.Int64, 'c': pl.Int64}) result_df = df.with_columns( pl.col('a').map_batches(batch_f) ) print(result_df)
该方式直接返回包含目标列的DataFrame,无需额外拆列步骤,适合大规模数据场景。
内容的提问来源于stack exchange,提问作者Des1303
相关产品推荐
相关产品推荐

