You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:基于unigram精确匹配规则过滤bigram词组列表

问题说明

现有一个存储bigram(二元词组)的列表,需要筛除所有不满足「至少有一个词与unigram(一元词)列表中的任意元素精确匹配」条件的二元词组。

两个输入列表如下

bigram_list = ['computer vision', 'data excellence', 'data visualization']
unigram_list = ['excel', 'tableau', 'visio', 'visualization']

预期输出目标

cleaned_bigrams = ['data visualization']

此前尝试过的无效方案

  • 适配Stack Overflow上《Removing separate list of items from another list in Python 3.x》的方案,未成功
  • 参考《Get rid of unigrams in a list if contained within bigrams or trigrams python》的方案实现,未达到预期效果
  • 适配此前提问的《Create new boolean fields based on specific bigrams appearing in a tokenized pandas dataframe》中的实现思路,未成功
解决方法

核心要抓住两个关键点:一是必须把每个二元词组拆成独立的单个词,二是匹配必须是精确全词匹配,不能做子串模糊匹配。
为了提升查找效率,先把unigram列表转成集合结构,之后遍历所有bigram做筛选即可:

# 将unigram转为集合,精确匹配查找效率远高于列表
unigram_pool = set(unigram_list)

cleaned_bigrams = []
for bigram in bigram_list:
    # 按空格拆分二元词组得到两个独立单词
    words = bigram.split()
    # 只要任意一个单词精确存在于unigram池中就保留
    if any(word in unigram_pool for word in words):
        cleaned_bigrams.append(bigram)

如果喜欢更简洁的写法,也可以用列表推导式实现和上面完全一致的逻辑:

unigram_pool = set(unigram_list)
cleaned_bigrams = [b for b in bigram_list if any(w in unigram_pool for w in b.split())]

之前方案不生效的原因

  • 列表元素移除类方案的逻辑是整串匹配,不会把bigram拆成单个词做比对,自然无法命中符合条件的词组
  • bigram/trigram包含unigram的清理方案大多用子串匹配逻辑,比如excellence包含excel这个子串,会被误判为匹配,不符合精确匹配的要求
  • pandas字段判断的方案面向整token列的匹配场景,没有拆分bigram做逐词精确校验的逻辑,直接套用无法得到正确结果

运行上述代码后得到的cleaned_bigrams正好是['data visualization'],和预期输出完全一致。

内容的提问来源于stack exchange,提问作者Cary Cox

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 06:17:02