You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中多结构体列复用字段表达式的实现方法及结果差异原因咨询

Polars中多结构体列复用字段表达式的实现方法及结果差异原因咨询

我最近在使用Polars处理带结构体的DataFrame时遇到了困惑,先给大家看看我的初始数据:

import polars as pl
import pandas as pd

df = pl.DataFrame({
    "dt":pd.date_range("2024-12-01", periods=4, freq="D"),
    "A": {"value":[10, 10, 20, 30], "multiple":[1,2,0,0]},
    "B": {"value":[1, 2, 3, 4], "multiple":[0,5,5,0]},
})

我本来想批量处理除dt之外的所有结构体列,给每个结构体都添加result(value乘multiple)和result2(result除以2)这两个字段,于是写了下面这段代码:

result = df.with_columns(
    pl.exclude("dt").struct.with_fields(
        result=pl.field("value") * pl.field("multiple"),
    ).struct.with_fields(
        result2= pl.field("result") / 2
    ),
)

print(
    result.select(
        pl.exclude("dt").struct.unnest().name.prefix("<dont know how to determine col name>"),
    )
)

但运行后得到的schema完全不是我预期的样子,我原本以为这段代码会给每一个排除dt的列单独添加对应的字段,结果却不是这么回事。

我想要的其实是下面这种效果,但这段代码里针对A和B的处理逻辑完全重复了,我不想这么写,也不想用Python列表推导式(比如*[pl.col(c).struct.with_fields(...) for c in wanted_cols])来实现:

result = df.with_columns(
    pl.col("A").struct.with_fields(
        result=pl.field("value") * pl.field("multiple"),
    ).struct.with_fields(
        result2= pl.field("result") / 2
    ),
    pl.col("B").struct.with_fields(
        result=pl.field("value") * pl.field("multiple"),
    ).struct.with_fields(
        result2= pl.field("result") / 2
    )
)

print(
    result.select(
        pl.col("A").struct.unnest().name.prefix("A_"),
        pl.col("B").struct.unnest().name.prefix("B_"),
    )
)

这段代码能输出我想要的结果:

shape: (4, 8)
┌─────────┬────────────┬──────────┬───────────┬─────────┬────────────┬──────────┬───────────┐
│ A_value ┆ A_multiple ┆ A_result ┆ A_result2 ┆ B_value ┆ B_multiple ┆ B_result ┆ B_result2 │
│ ---     ┆ ---        ┆ ---      ┆ ---       ┆ ---     ┆ ---        ┆ ---      ┆ ---       │
│ i64     ┆ i64        ┆ i64      ┆ f64       ┆ i64     ┆ i64        ┆ i64      ┆ f64       │
╞═════════╪════════════╪══════════╪═══════════╪═════════╪════════════╪══════════╪═══════════╡
│ 10      ┆ 1          ┆ 10       ┆ 5.0       ┆ 1       ┆ 0          ┆ 0        ┆ 0.0       │
│ 10      ┆ 2          ┆ 20       ┆ 10.0      ┆ 2       ┆ 5          ┆ 10       ┆ 5.0       │
│ 20      ┆ 0          ┆ 0        ┆ 0.0       ┆ 3       ┆ 5          ┆ 15       ┆ 7.5       │
│ 30      ┆ 0          ┆ 0        ┆ 0.0       ┆ 4       ┆ 0          ┆ 0        ┆ 0.0       │
└─────────┴────────────┴──────────┴───────────┴─────────┴────────────┴──────────┴───────────┘

我现在的疑问是:为什么第一种实现方式得到的结构体schema和第二种不一样呢?

另外我还尝试过把数据改成扁平结构(没有结构体,列名是A_value、A_multiple、B_value、B_multiple这类),然后用下面的代码处理:

df.with_columns(
    (pl.selectors.ends_with("_value") * pl.selectors.ends_with("_multiple")).name.map(lambda c: c.split("_")[0] + "_result")
)

但这种同时使用不同正则选择器的表达式目前Polars还不支持。

备注:内容来源于stack exchange,提问作者Jan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 17:34:30