You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅删除DataFrame文本列中连续出现的指定两个单词?

连续双词删除实现方案

以下两种方案都可以满足你的需求,可根据实际场景选择:

方法1:正则快速替换(适合固定词组的简单场景)

直接用正则匹配精准定位连续出现的指定双词,替换后清理多余空格即可,代码量最小:

import pandas as pd

# 定义要删除的连续词组,\b 为单词边界符,避免误匹配包含目标词片段的长单词
target_pattern = r'\bboy live\b'
# 执行替换并清理多余空格
df['cleaned_text'] = df['text'].str.replace(target_pattern, '', regex=True)\
                               .str.strip()\
                               .str.replace(r'\s{2,}', ' ', regex=True)

处理后得到的cleaned_text列就是你要的结果,你的示例数据运行后输出和预期完全一致。

方法2:分词滑动匹配(适合多组连续词组删除的复杂场景)

如果你后续需要批量删除多组不同的连续双词,或者要结合现有的nltk停用词过滤逻辑一起使用,可以用自定义函数实现:

import pandas as pd

def clean_text(text, target_pairs, stopwords=None):
    words = text.split()
    result = []
    idx = 0
    word_count = len(words)
    while idx < word_count:
        # 先判断是否是要删除的连续双词
        if idx < word_count -1 and (words[idx], words[idx+1]) in target_pairs:
            idx += 2
            continue
        # 可选:叠加原有停用词过滤逻辑
        if stopwords and words[idx] in stopwords:
            idx +=1
            continue
        result.append(words[idx])
        idx +=1
    return ' '.join(result)

# 配置要删除的连续双词组,支持同时配置多组
target_pairs = {('boy', 'live')}
# 如果你要继续用原来的nltk停用词,把停用词集合传进去即可
# from nltk.corpus import stopwords
# stop_words = set(stopwords.words('english'))
# df['cleaned_text'] = df['text'].apply(lambda x: clean_text(x, target_pairs, stop_words))

# 不需要停用词过滤的话用这行
df['cleaned_text'] = df['text'].apply(lambda x: clean_text(x, target_pairs))

内容的提问来源于stack exchange,提问作者R.A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 11:36:07