You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何无需循环过滤Polars DataFrame中与另一数据框字符串部分匹配的行?

如何在Polars中批量过滤包含任意匹配模式的行(无需循环)

需求说明

需要删除主DataFrame中,filename列包含过滤DataFrame中任意字符串的所有行,替代逐行循环的低效实现,同时对齐SQL中的批量过滤逻辑。

方法1:构建正则表达式批量匹配

将过滤条件中的唯一值拼接成正则表达式的OR模式,通过一次str.contains完成批量匹配,性能最优:

import polars as pl

df_data = pl.read_csv("data.csv")
df_filter = pl.read_csv("filter.csv")

# 提取去重后的过滤词,拼接为正则匹配模式
filter_terms = df_filter.select(pl.col("skip").unique()).to_series().to_list()
match_pattern = "|".join(filter_terms)

# 过滤掉包含任意过滤词的行
df_cleaned = df_data.filter(~pl.col("filename").str.contains(match_pattern))

print("data after:\n", df_cleaned)

方法2:模拟SQL的EXISTS逻辑(反连接实现)

通过交叉连接找到所有匹配的行,再用反连接保留不匹配的行,完全对齐PostgreSQL中WHERE EXISTS的逻辑:

import polars as pl

df_data = pl.read_csv("data.csv")
# 先对过滤表去重,避免重复匹配
df_filter_unique = pl.read_csv("filter.csv").unique(subset="skip")

# 找到所有需要删除的匹配行
matching_rows = df_data.join(
    df_filter_unique,
    how="cross"
).filter(
    pl.col("filename").str.contains(pl.col("skip"))
).select(df_data.columns).unique()

# 反连接保留未匹配的行
df_cleaned = df_data.join(matching_rows, on=df_data.columns, how="anti")

print("data after:\n", df_cleaned)

与PostgreSQL实现的逻辑映射

  • PostgreSQL中通过UPDATE ... WHERE EXISTS标记需要删除的行再执行删除,对应Polars的反连接方法:先定位匹配行,再排除这些行。
  • PostgreSQL中LIKE s.skip的批量匹配,对应Polars中用正则|拼接多个匹配模式的方式,本质都是将多个匹配条件合并为一次批量检查。

输出结果

两种方法均会得到正确的过滤结果:

data after:
 shape: (4, 2)
┌────────────┬─────────┐
│ filename   ┆ col2    │
│ ---        ┆ ---     │
│ str        ┆ str     │
├────────────┼─────────┤
│ keep.txt   ┆ bar     │
│ keep2.txt  ┆ zoom    │
│ file3.txt  ┆ custom3 │
│ file4.txt  ┆ custom5 │
└────────────┴─────────┘

内容的提问来源于stack exchange,提问作者JL Peyret

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 15:13:10