如何仅删除DataFrame文本列中连续出现的指定两个单词?
连续双词删除实现方案
以下两种方案都可以满足你的需求,可根据实际场景选择:
方法1:正则快速替换(适合固定词组的简单场景)
直接用正则匹配精准定位连续出现的指定双词,替换后清理多余空格即可,代码量最小:
import pandas as pd # 定义要删除的连续词组,\b 为单词边界符,避免误匹配包含目标词片段的长单词 target_pattern = r'\bboy live\b' # 执行替换并清理多余空格 df['cleaned_text'] = df['text'].str.replace(target_pattern, '', regex=True)\ .str.strip()\ .str.replace(r'\s{2,}', ' ', regex=True)
处理后得到的cleaned_text列就是你要的结果,你的示例数据运行后输出和预期完全一致。
方法2:分词滑动匹配(适合多组连续词组删除的复杂场景)
如果你后续需要批量删除多组不同的连续双词,或者要结合现有的nltk停用词过滤逻辑一起使用,可以用自定义函数实现:
import pandas as pd def clean_text(text, target_pairs, stopwords=None): words = text.split() result = [] idx = 0 word_count = len(words) while idx < word_count: # 先判断是否是要删除的连续双词 if idx < word_count -1 and (words[idx], words[idx+1]) in target_pairs: idx += 2 continue # 可选:叠加原有停用词过滤逻辑 if stopwords and words[idx] in stopwords: idx +=1 continue result.append(words[idx]) idx +=1 return ' '.join(result) # 配置要删除的连续双词组,支持同时配置多组 target_pairs = {('boy', 'live')} # 如果你要继续用原来的nltk停用词,把停用词集合传进去即可 # from nltk.corpus import stopwords # stop_words = set(stopwords.words('english')) # df['cleaned_text'] = df['text'].apply(lambda x: clean_text(x, target_pairs, stop_words)) # 不需要停用词过滤的话用这行 df['cleaned_text'] = df['text'].apply(lambda x: clean_text(x, target_pairs))
内容的提问来源于stack exchange,提问作者R.A
相关产品推荐
相关产品推荐

