如何在Pandas DataFrame列中批量移除指定词汇?
解决Pandas文本列批量移除指定词汇的问题
原代码的问题分析
你的函数未生效核心原因有两个:
- 错误遍历字符串单个字符:
for i in t会逐个遍历字符串的字符,而非拆分后的单词,根本无法匹配目标词汇 - 忽略字符串不可变性:
t.replace(i, '')返回新字符串但未赋值给变量,且函数没有返回处理后的结果
方法一:自定义过滤函数(灵活易读)
通过「拆分单词-过滤目标-重新拼接」的逻辑编写函数:
def remove_target_words(text): # 用集合存储目标词汇,查询效率更高 target_words = {'Livre', 'Chapitre', 'Titre', 'Chapter', 'Article'} # 拆分文本为单词列表,过滤掉目标词汇后重新拼接 filtered_words = [word for word in text.split() if word not in target_words] return ' '.join(filtered_words)
测试示例:
st = 'this is Livre and Chapitre and Titre and Chapter and Article' print(remove_target_words(st)) # 输出: this is and and and and
批量应用到DataFrame的content列:
df['content'] = df['content'].apply(remove_target_words)
方法二:Pandas内置正则替换(高效适配大数据)
利用Pandas的str.replace结合正则匹配整个单词,处理效率更优:
import re target_words = ['Livre', 'Chapitre', 'Titre', 'Chapter', 'Article'] # 构建正则模式,\b 匹配单词边界,re.escape处理词汇中的特殊字符 pattern = r'\b(' + '|'.join(re.escape(word) for word in target_words) + r')\b' # 替换目标词汇,可选str.strip()移除首尾多余空格 df['content'] = df['content'].str.replace(pattern, '', regex=True)
若需清理替换后产生的连续空格,可追加一步:
df['content'] = df['content'].str.replace(r'\s+', ' ', regex=True)
内容的提问来源于stack exchange,提问作者ForeverLearner
相关产品推荐
相关产品推荐

