You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars DataFrame中用列表列过滤另一列表列?

Polars:过滤列表列中已存在于另一列表的元素

方法1:逐行遍历过滤(保持原顺序)

用pl.map_elements结合列表推导式,直接对每行的两个列表进行过滤,严格保留popular_items中的元素顺序:

import polars as pl

df = pl.select(user_id=1, items=[1, 2, 3, 4], popular_items=[3, 4, 5, 6])

df = df.with_columns(
    suggested=pl.map_elements(
        lambda user_items, pop_items: [item for item in pop_items if item not in user_items],
        args=["items", "popular_items"],
        return_dtype=pl.List(pl.Int64)
    )
)

print(df)

输出结果:

┌─────────────┬─────────────┬───────────────┬───────────┐
│ user_id     ┆ items       ┆ popular_items ┆ suggested │
│ ---         ┆ ---         ┆ ---           ┆ ---       │
│ i64         ┆ list[i64]   ┆ list[i64]     ┆ list[i64] │
╞═════════════╪═════════════╪═══════════════╪═══════════╡
│ 1           ┆ [1, 2, 3, 4]┆ [3, 4, 5, 6]  ┆ [5, 6]    │
└─────────────┴─────────────┴───────────────┴───────────┘

该方法逻辑直观,适合小规模数据集;若处理超大规模数据,建议使用向量化方法。

方法2:内置集合差集(高效但可能打乱顺序)

Polars提供list.set_difference函数,直接计算两个列表的集合差集,效率更高,但结果是无序的(集合操作不保留顺序):

df = df.with_columns(
    suggested=pl.col("popular_items").list.set_difference(pl.col("items"))
)

输出可能类似[6,5],如果业务不要求保持原popular_items的顺序,这个方法是最优选择。

方法3:向量化展开过滤(大数据量友好,保持顺序)

通过explode展开列表,过滤掉存在于items中的元素后重新聚合,全程使用Polars向量化操作,适合处理大规模数据集,且能严格保留原顺序:

df_result = (
    df.with_row_index("row_idx")
      .explode("popular_items")
      .filter(~pl.col("popular_items").is_in(pl.col("items")))
      .group_by("row_idx", maintain_order=True)
      .agg(
          pl.col("user_id").first(),
          pl.col("items").first(),
          pl.col("popular_items").alias("suggested")
      )
      .drop("row_idx")
)

print(df_result)

该方法避免了逐行Python函数调用,性能更优,同时通过maintain_order=True保证聚合后的顺序与原数据一致。

内容的提问来源于stack exchange,提问作者D_usv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 01:40:21