Polars中多结构体列复用字段表达式的实现方法及结果差异原因咨询
Polars中多结构体列复用字段表达式的实现方法及结果差异原因咨询
我最近在使用Polars处理带结构体的DataFrame时遇到了困惑,先给大家看看我的初始数据:
import polars as pl import pandas as pd df = pl.DataFrame({ "dt":pd.date_range("2024-12-01", periods=4, freq="D"), "A": {"value":[10, 10, 20, 30], "multiple":[1,2,0,0]}, "B": {"value":[1, 2, 3, 4], "multiple":[0,5,5,0]}, })
我本来想批量处理除dt之外的所有结构体列,给每个结构体都添加result(value乘multiple)和result2(result除以2)这两个字段,于是写了下面这段代码:
result = df.with_columns( pl.exclude("dt").struct.with_fields( result=pl.field("value") * pl.field("multiple"), ).struct.with_fields( result2= pl.field("result") / 2 ), ) print( result.select( pl.exclude("dt").struct.unnest().name.prefix("<dont know how to determine col name>"), ) )
但运行后得到的schema完全不是我预期的样子,我原本以为这段代码会给每一个排除dt的列单独添加对应的字段,结果却不是这么回事。
我想要的其实是下面这种效果,但这段代码里针对A和B的处理逻辑完全重复了,我不想这么写,也不想用Python列表推导式(比如*[pl.col(c).struct.with_fields(...) for c in wanted_cols])来实现:
result = df.with_columns( pl.col("A").struct.with_fields( result=pl.field("value") * pl.field("multiple"), ).struct.with_fields( result2= pl.field("result") / 2 ), pl.col("B").struct.with_fields( result=pl.field("value") * pl.field("multiple"), ).struct.with_fields( result2= pl.field("result") / 2 ) ) print( result.select( pl.col("A").struct.unnest().name.prefix("A_"), pl.col("B").struct.unnest().name.prefix("B_"), ) )
这段代码能输出我想要的结果:
shape: (4, 8) ┌─────────┬────────────┬──────────┬───────────┬─────────┬────────────┬──────────┬───────────┐ │ A_value ┆ A_multiple ┆ A_result ┆ A_result2 ┆ B_value ┆ B_multiple ┆ B_result ┆ B_result2 │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ i64 ┆ f64 ┆ i64 ┆ i64 ┆ i64 ┆ f64 │ ╞═════════╪════════════╪══════════╪═══════════╪═════════╪════════════╪══════════╪═══════════╡ │ 10 ┆ 1 ┆ 10 ┆ 5.0 ┆ 1 ┆ 0 ┆ 0 ┆ 0.0 │ │ 10 ┆ 2 ┆ 20 ┆ 10.0 ┆ 2 ┆ 5 ┆ 10 ┆ 5.0 │ │ 20 ┆ 0 ┆ 0 ┆ 0.0 ┆ 3 ┆ 5 ┆ 15 ┆ 7.5 │ │ 30 ┆ 0 ┆ 0 ┆ 0.0 ┆ 4 ┆ 0 ┆ 0 ┆ 0.0 │ └─────────┴────────────┴──────────┴───────────┴─────────┴────────────┴──────────┴───────────┘
我现在的疑问是:为什么第一种实现方式得到的结构体schema和第二种不一样呢?
另外我还尝试过把数据改成扁平结构(没有结构体,列名是A_value、A_multiple、B_value、B_multiple这类),然后用下面的代码处理:
df.with_columns( (pl.selectors.ends_with("_value") * pl.selectors.ends_with("_multiple")).name.map(lambda c: c.split("_")[0] + "_result") )
但这种同时使用不同正则选择器的表达式目前Polars还不支持。
备注:内容来源于stack exchange,提问作者Jan
相关产品推荐
相关产品推荐

