You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas文本替换添加条件:排除前3词含no/none的场景

带规则的Pandas文本替换解决方案

嘿,我来帮你搞定这个带条件的文本替换需求!你要实现的是:在文本中搜索替换指定术语,但如果目标术语的前3个词范围内出现了"no"或"none",就跳过这个替换——就像你例子里ID=2的条目,因为"random text"前面有"no",所以不能被替换对吧?

下面是具体的实现步骤和代码:

1. 准备示例数据

首先我们先定义你给出的文本DataFrame,再模拟一个包含替换术语的DataFrame(你可以根据实际情况修改术语和替换值):

import pandas as pd
import re

# 你的文本数据
text_data = {'ID': [1, 2, 3], 'Text': ['here is some random text', 'no such random text, none here', 'more sample content']}
text_df = pd.DataFrame(text_data)

# 术语替换表(自定义你的术语和替换值)
terms_df = pd.DataFrame({
    'Term': ['random text', 'sample content'],
    'Replacement': ['REPLACED_TEXT', 'REPLACED_CONTENT']
})

2. 编写带规则的替换函数

核心逻辑是:对每一段文本,遍历所有要替换的术语,找到术语的所有出现位置,然后检查该位置前3个词是否包含禁词("no"或"none"),只有不包含的时候才执行替换。

def replace_with_rule(text, terms):
    # 定义禁词集合,不区分大小写(如果需要区分可以去掉lower())
    forbidden_words = {'no', 'none'}
    
    # 遍历每个替换术语
    for _, term_row in terms.iterrows():
        target_term = term_row['Term']
        replacement = term_row['Replacement']
        
        # 找到文本中所有术语的匹配位置(用re.escape避免术语里的正则特殊字符干扰)
        matches = list(re.finditer(re.escape(target_term), text))
        
        # 从后往前替换,避免替换后文本长度变化导致的位置偏移问题
        for match in reversed(matches):
            match_start = match.start()
            # 获取术语出现位置之前的文本
            preceding_text = text[:match_start]
            # 分割成词,取最后3个(如果不足3个就取全部)
            preceding_words = preceding_text.split()[-3:]
            
            # 检查前3个词里是否有禁词
            if any(word.lower() in forbidden_words for word in preceding_words):
                # 有禁词,跳过这个替换
                continue
            
            # 没有禁词,执行替换
            text = text[:match_start] + replacement + text[match.end():]
    
    return text

3. 应用函数到DataFrame

把上面的函数应用到文本列,生成处理后的新列:

text_df['Processed_Text'] = text_df['Text'].apply(lambda x: replace_with_rule(x, terms_df))

4. 查看结果

运行后你会得到这样的输出:

IDTextProcessed_Text
1here is some random texthere is some REPLACED_TEXT
2no such random text, none hereno such random text, none here
3more sample contentmore REPLACED_CONTENT

可以看到:

  • ID=1的"random text"前3个词是"here is some",没有禁词,成功替换
  • ID=2的"random text"前3个词包含"no",跳过替换
  • ID=3的"sample content"前只有1个词"more",没有禁词,成功替换

关键细节说明

  • re.escape(target_term):用来处理术语中可能包含的正则特殊字符(比如"."、"*"),避免它们被当成正则语法解析
  • 从后往前替换:如果从前往后替换,替换后的文本长度变化会导致后续匹配的位置偏移,从后往前就不会有这个问题
  • 不区分大小写检查:代码里用了word.lower(),如果你的需求是只匹配小写的"no"/"none",可以去掉这个转换

内容的提问来源于stack exchange,提问作者shbfy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:26:37