Polars中如何基于selector列从结构体列表列提取匹配字段?
解决方案
方法1:使用list.filter + 字段提取(最简洁)
直接通过list.filter筛选出s列中x字段与当前行selector匹配的结构体,再提取其中的y字段,最后取第一个匹配结果(假设每行仅存在一个匹配项):
import polars as pl # 示例数据 df = pl.DataFrame({ "id": [1, 2], "s": [ [{"x": "a", "y": 10}, {"x": "b", "y": 20}], [{"x": "c", "y": 30}, {"x": "d", "y": 40}] ], "selector": ["b", "c"] }) # 提取匹配的y字段 result = df.with_columns( matched_y=pl.col("s") .list.filter(pl.element().struct.field("x") == pl.col("selector")) .list.struct.field("y") .list.get(0) ) print(result)
输出结果:
shape: (2, 4) ┌─────┬──────────────────────────────────┬───────────┬───────────┐ │ id ┆ s ┆ selector ┆ matched_y │ │ --- ┆ --- ┆ --- ┆ --- │ │ i64 ┆ list[struct[2]] ┆ str ┆ i64 │ ╞═════╪══════════════════════════════════╪═══════════╪═══════════╡ │ 1 ┆ [{"a", 10}, {"b", 20}] ┆ b ┆ 20 │ │ 2 ┆ [{"c", 30}, {"d", 40}] ┆ c ┆ 30 │ └─────┴──────────────────────────────────┴───────────┴───────────┘
方法2:修正list.eval的用法
如果坚持用list.eval,需要通过pl.col("selector").first()来引用当前行的selector值(因为list.eval的上下文是列表内部,直接用命名列会报错),之后过滤并提取字段:
result = df.with_columns( matched_y=pl.col("s") .list.eval( pl.when(pl.element().struct.field("x") == pl.col("selector").first()) .then(pl.element().struct.field("y")) ) .list.flatten() # 过滤掉None值 .list.get(0) )
为什么之前的list.eval会报错?
list.eval的执行上下文是列表的每个元素,默认无法直接访问外部的命名列(比如selector),必须通过pl.col("selector").first()明确获取当前行的selector值,否则会触发ComputeError。
内容的提问来源于stack exchange,提问作者bzm3r
相关产品推荐
相关产品推荐

