You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检测DataFrame的text列含指定关键词并替换内容?

解决DataFrame文本含关键词替换问题

场景与需求

给定示例DataFrame:

df = [{'id': 1, 'text': 'text contains ok words'}, {'id':2, 'text':'text contains word apple'}, {'id':3, 'text':'text contains words ok'}]

关键词列表:

keywords = ['apple', 'orange', 'lime']

需求:检查text列内容,若包含列表中任意关键词,则将该单元格内容替换为'disconsider this case'。

你尝试代码的问题点

自定义函数错误

你写的remove_keywords函数存在两处问题:

  1. 变量名错误:keyword应为keywords
  2. 逻辑错误:遍历逻辑颠倒,应该检查关键词是否在输入文本中,而非反过来

分词遍历错误

第二种方法的问题:

  1. 'apple' in df['text']是检查字符串是否在整个Series中,而非当前行的分词列表里
  2. 循环中无法直接用return修改DataFrame内容

正确解决办法

方法1:使用pandas内置字符串方法(推荐)

利用str.contains直接匹配关键词,无需手动分词,效率最高:

import pandas as pd

# 转换为DataFrame
df = pd.DataFrame(df)
keywords = ['apple', 'orange', 'lime']

# 构建正则表达式,\b确保匹配完整单词(避免'applepie'这类部分匹配)
pattern = r'\b(' + '|'.join(keywords) + r')\b'

# 替换符合条件的行
df.loc[df['text'].str.contains(pattern, case=False), 'text'] = 'disconsider this case'

方法2:修复自定义函数

如果偏好自定义函数,修正逻辑后实现:

import pandas as pd

df = pd.DataFrame(df)
keywords = ['apple', 'orange', 'lime']

def remove_keywords(text):
    # 检查文本中是否包含任意关键词,处理首尾单词的情况
    text_with_spaces = f' {text} '
    return 'disconsider this case' if any(f' {kw} ' in text_with_spaces for kw in keywords) else text

df['text'] = df['text'].apply(remove_keywords)

方法3:分词后匹配(适合复杂文本)

如果需要先分词再精准匹配单词:

import pandas as pd
import nltk
nltk.download('punkt')

df = pd.DataFrame(df)
keywords = ['apple', 'orange', 'lime']
keyword_set = set(keywords)

def check_keywords(text):
    # 分词并转小写,避免大小写干扰
    words = nltk.word_tokenize(text.lower())
    # 检查分词结果与关键词集合是否有交集
    return 'disconsider this case' if keyword_set & set(words) else text

df['text'] = df['text'].apply(check_keywords)

内容的提问来源于stack exchange,提问作者G Nova

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 08:01:05