You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中检测数据集行是否包含短语列表的全部单词

解决方案

核心思路

把每个目标短语拆解为小写单词集合,再将每行搜索文本预处理为去除标点的小写单词集合,通过判断「是否存在某个短语的单词集合是当前文本单词集合的子集」来标记匹配结果。这种方式既不关心单词顺序,也允许文本包含其他内容。

完整代码

import pandas as pd
import re

# 示例数据集
text = [('how to screenshot on mac', 0),
         ('how to take screenshot?', 0),
         ('how to take screenshot on windows', 0),
         ('when is christmas', 0),
         ('how many days until christmas', 0),
        ('how many weeks until christmas', 0),
        ('how much is the new google pixel 8', 0),
        ('which google pixel versions are available', 0),
        ('how do I do google search on my pixel phone 7a', 0)]
labels = ['Search','Random_Column']
df = pd.DataFrame.from_records(text, columns=labels)

# 目标短语列表
phrases = ['mac screenshot', 'days until christmas', 'google pixel 7a']

# 1. 将短语转换为小写单词集合
phrase_sets = [set(phrase.lower().split()) for phrase in phrases]

# 2. 定义匹配判断函数
def check_match(search_text):
    # 预处理:转小写、去除标点、拆分为单词集合
    processed_text = re.sub(r'[^\w\s]', '', search_text.lower()).split()
    text_set = set(processed_text)
    # 检查是否有任意短语的单词集合是当前文本集合的子集
    return any(phrase.issubset(text_set) for phrase in phrase_sets)

# 3. 生成Match列
df['Match'] = df['Search'].apply(check_match)

print(df)

输出结果

Search  Random_Column  Match
0                  how to screenshot on mac              0   True
1                     how to take screenshot?              0  False
2          how to take screenshot on windows              0  False
3                          when is christmas              0  False
4              how many days until christmas              0   True
5             how many weeks until christmas              0  False
6         how much is the new google pixel 8              0  False
7  which google pixel versions are available              0  False
8  how do I do google search on my pixel phone 7a              0   True

关键细节说明

  • 标点处理:用re.sub(r'[^\w\s]', '', text)去除所有非字母、数字、空格的字符,避免像screenshot?这类带标点的单词无法匹配screenshot。
  • 大小写统一:所有文本转小写,避免大小写差异导致的匹配失败(比如Mac和mac视为同一个单词)。
  • 子集判断:issubset()方法直接判断短语的所有单词是否都存在于当前文本中,完美满足「不关心顺序、允许其他内容」的需求。

内容的提问来源于stack exchange,提问作者Maria

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 23:33:11