如何用Pandas在DataFrame列中匹配列表多词并生成新列
Pandas匹配多关键词并生成匹配结果列(匹配数≥3时)
需求说明
给定关键词列表:
keyword_list = ['motorcycle love hobby ', 'bike love me', 'cycle', 'dirtbike cycle motorbike ']
需要在Pandas DataFrame的指定列中匹配关键词,当某行匹配到3个及以上独立词语时,在新列中存储所有匹配到的词语;未达3个则返回空值。
实现步骤与代码
1. 预处理关键词,生成去重的词语集合
先把原关键词列表里的所有条目拆成单个词语,去除空格和重复项,方便后续匹配:
import pandas as pd # 原关键词列表 keyword_list = ['motorcycle love hobby ', 'bike love me', 'cycle', 'dirtbike cycle motorbike '] # 拆分并清洗关键词,生成去重的词语集合 keywords = set() for item in keyword_list: # 拆分词语、去除空字符串(处理末尾空格)、加入集合去重 words = [w.strip() for w in item.split() if w.strip()] keywords.update(words)
2. 构造示例DataFrame
这里模拟业务数据,你可以替换成自己的真实数据:
# 示例数据 data = { 'text': [ 'I love my motorcycle and hobby, also ride a cycle', 'Just a bike for me', 'Dirtbike, motorbike and cycle are my favorite, love them', 'Only ride a cycle sometimes' ] } df = pd.DataFrame(data)
3. 定义匹配函数并生成新列
编写函数处理每行文本,提取匹配的词语,判断数量是否达标:
def match_keywords(text): # 拆分当前文本为词语,转小写(实现大小写不敏感匹配,按需可移除) text_words = [w.strip().lower() for w in text.split() if w.strip()] # 匹配关键词 matched = [word for word in keywords if word.lower() in text_words] # 匹配数≥3则返回列表,否则返回空值 return matched if len(matched) >= 3 else None # 应用函数生成新列'matched_keywords' df['matched_keywords'] = df['text'].apply(match_keywords)
4. 最终结果
运行后得到的DataFrame效果如下:
| text | matched_keywords |
|---|---|
| I love my motorcycle and hobby, also ride a cycle | ['love', 'motorcycle', 'hobby', 'cycle'] |
| Just a bike for me | None |
| Dirtbike, motorbike and cycle are my favorite, love them | ['dirtbike', 'motorbike', 'cycle', 'love'] |
| Only ride a cycle sometimes | None |
注意事项
- 若需要大小写敏感匹配,去掉代码中
.lower()的处理即可; - 如果文本中有标点符号,可先通过
re.sub(r'[^\w\s]', '', text)去除标点后再拆分词语; - 关键词列表中的重复词语会被自动去重,避免重复统计。
内容的提问来源于stack exchange,提问作者rudra
相关产品推荐
相关产品推荐

