Pandas文本替换添加条件:排除前3词含no/none的场景
带规则的Pandas文本替换解决方案
嘿,我来帮你搞定这个带条件的文本替换需求!你要实现的是:在文本中搜索替换指定术语,但如果目标术语的前3个词范围内出现了"no"或"none",就跳过这个替换——就像你例子里ID=2的条目,因为"random text"前面有"no",所以不能被替换对吧?
下面是具体的实现步骤和代码:
1. 准备示例数据
首先我们先定义你给出的文本DataFrame,再模拟一个包含替换术语的DataFrame(你可以根据实际情况修改术语和替换值):
import pandas as pd import re # 你的文本数据 text_data = {'ID': [1, 2, 3], 'Text': ['here is some random text', 'no such random text, none here', 'more sample content']} text_df = pd.DataFrame(text_data) # 术语替换表(自定义你的术语和替换值) terms_df = pd.DataFrame({ 'Term': ['random text', 'sample content'], 'Replacement': ['REPLACED_TEXT', 'REPLACED_CONTENT'] })
2. 编写带规则的替换函数
核心逻辑是:对每一段文本,遍历所有要替换的术语,找到术语的所有出现位置,然后检查该位置前3个词是否包含禁词("no"或"none"),只有不包含的时候才执行替换。
def replace_with_rule(text, terms): # 定义禁词集合,不区分大小写(如果需要区分可以去掉lower()) forbidden_words = {'no', 'none'} # 遍历每个替换术语 for _, term_row in terms.iterrows(): target_term = term_row['Term'] replacement = term_row['Replacement'] # 找到文本中所有术语的匹配位置(用re.escape避免术语里的正则特殊字符干扰) matches = list(re.finditer(re.escape(target_term), text)) # 从后往前替换,避免替换后文本长度变化导致的位置偏移问题 for match in reversed(matches): match_start = match.start() # 获取术语出现位置之前的文本 preceding_text = text[:match_start] # 分割成词,取最后3个(如果不足3个就取全部) preceding_words = preceding_text.split()[-3:] # 检查前3个词里是否有禁词 if any(word.lower() in forbidden_words for word in preceding_words): # 有禁词,跳过这个替换 continue # 没有禁词,执行替换 text = text[:match_start] + replacement + text[match.end():] return text
3. 应用函数到DataFrame
把上面的函数应用到文本列,生成处理后的新列:
text_df['Processed_Text'] = text_df['Text'].apply(lambda x: replace_with_rule(x, terms_df))
4. 查看结果
运行后你会得到这样的输出:
| ID | Text | Processed_Text |
|---|---|---|
| 1 | here is some random text | here is some REPLACED_TEXT |
| 2 | no such random text, none here | no such random text, none here |
| 3 | more sample content | more REPLACED_CONTENT |
可以看到:
- ID=1的"random text"前3个词是"here is some",没有禁词,成功替换
- ID=2的"random text"前3个词包含"no",跳过替换
- ID=3的"sample content"前只有1个词"more",没有禁词,成功替换
关键细节说明
re.escape(target_term):用来处理术语中可能包含的正则特殊字符(比如"."、"*"),避免它们被当成正则语法解析- 从后往前替换:如果从前往后替换,替换后的文本长度变化会导致后续匹配的位置偏移,从后往前就不会有这个问题
- 不区分大小写检查:代码里用了
word.lower(),如果你的需求是只匹配小写的"no"/"none",可以去掉这个转换
内容的提问来源于stack exchange,提问作者shbfy
相关产品推荐
相关产品推荐

