You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中如何按行对两列列表执行search_sorted操作?

如何在Polars DataFrame中按行对两列列表执行search_sorted操作?

我尝试在Polars DataFrame的a列与b列的列表之间执行search_sorted排序搜索操作:

import polars as pl

df = pl.DataFrame(
    {
        "a": [[1, 2, 3], [8, 9]],
        "b": [[2], [10, 6]]
    }
)

print(df)

res = df.lazy().with_columns(
    pl.col("a").explode().search_sorted(pl.col("b").explode(), side="left").implode().alias("c")
)

print(res.collect())

得到的结果不符合预期:

shape: (2, 3)
┌───────────┬───────────┬───────────┐
│ a         ┆ b         ┆ c         │
│ ---       ┆ ---       ┆ ---       │
│ list[i64] ┆ list[i64] ┆ list[u32] │
╞═══════════╪═══════════╪═══════════╡
│ [1, 2, 3] ┆ [2]       ┆ [1, 5, 3] │
│ [8, 9]    ┆ [10, 6]   ┆ [1, 5, 3] │
└───────────┴───────────┴───────────┘

我期望按行计算,即第一行结果应为[1],第二行结果应为[2, 0]。

遇到的问题

  • 使用explode会将所有行的列表拼接成列级数据,无法实现行内计算。
  • 直接对列表列调用search_sorted会报错:
    exceptions.InvalidOperationError:不支持对`list[i64]`类型执行`search_sorted`操作
    
  • 尝试用list.eval但无法引用另一列的列表,报错:
    ComputeError:`list.eval`中不允许使用命名列;请考虑使用`element`或`col("")`
    

解决方案

可以利用Polars的list.eval结合行上下文的列引用,实现高效的行内列表search_sorted操作:

import polars as pl

df = pl.DataFrame(
    {
        "a": [[1, 2, 3], [8, 9]],
        "b": [[2], [10, 6]]
    }
)

res = df.with_columns(
    # 对b列的每个列表元素,在当前行的a列表中执行search_sorted
    pl.col("b")
        .list.eval(pl.col("a").first().search_sorted(pl.element(), side="left"), parallel=True)
        .alias("c")
)

print(res)

执行结果:

shape: (2, 3)
┌───────────┬───────────┬───────────┐
│ a         ┆ b         ┆ c         │
│ ---       ┆ ---       ┆ ---       │
│ list[i64] ┆ list[i64] ┆ list[u32] │
╞═══════════╪═══════════╪═══════════╡
│ [1, 2, 3] ┆ [2]       ┆ [1]       │
│ [8, 9]    ┆ [10, 6]   ┆ [2, 0]    │
└───────────┴───────────┴───────────┘

原理说明

  • pl.col("b").list.eval(...):遍历b列的每个列表元素。
  • pl.col("a").first():在list.eval的行上下文里,获取当前行的a列列表(每行仅一个a列表,first()即可拿到完整值)。
  • search_sorted(pl.element(), side="left"):对b列的每个元素(pl.element()指代当前遍历的元素),在当前行的a列表中执行左边界排序搜索。
  • parallel=True:开启并行计算提升效率,适合大数据量场景。

如果需要使用Lazy API,只需稍作调整:

res = df.lazy().with_columns(
    pl.col("b")
        .list.eval(pl.col("a").first().search_sorted(pl.element(), side="left"), parallel=True)
        .alias("c")
).collect()

内容的提问来源于stack exchange,提问作者The Unfun Cat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 09:37:06