如何在统计句子目标词出现次数时排除否定词影响?
解决方案
要实现排除否定词影响的单词计数,得替换掉简单的count方法,从文本标准化、否定词范围检测两个核心点入手,以下是具体实现:
核心思路
- 标准化文本:统一大小写、剥离标点、处理单复数,确保目标词匹配准确
- 定义否定词集合:直接包含
don't、not - 设置检测窗口:自定义目标词前后的单词范围(比如前后3个单词),范围内出现否定词就跳过计数
- 遍历匹配:逐个检查目标词的出现位置,判断是否符合计数条件
代码实现
import re def count_valid_words(target_words, text, window_size=3): # 定义需要排除的否定词 negation_words = {"don't", "not"} # 预处理文本:提取带缩写的单词,统一转小写 words = re.findall(r"\b[\w']+\b", text.lower()) count_result = {word: 0 for word in target_words} for target in target_words: # 处理单复数统一匹配(简单规则:去掉末尾的s,可按需扩展) normalized_target = target[:-1] if target.endswith('s') else target # 遍历所有单词的位置 for idx, word in enumerate(words): normalized_word = word[:-1] if word.endswith('s') else word if normalized_word == normalized_target: # 计算检测范围的上下边界,避免越界 check_start = max(0, idx - window_size) check_end = min(len(words), idx + window_size + 1) # 检查范围内是否存在否定词 has_negation = any(neg in words[check_start:check_end] for neg in negation_words) if not has_negation: count_result[target] += 1 return count_result # 测试用例 w = ["hello", "apple"] txt = "I love apples, apple are my favorite fruit. I don't really like apples if they are too mature. I do not like apples if they are immature either." final_counts = count_valid_words(w, txt, window_size=3) for word, count in final_counts.items(): print(f"{word}: {count}")
代码说明
- 文本预处理:用正则
\b[\w']+\b提取单词,保留don't这类缩写,同时转小写避免大小写干扰 - 单复数匹配:通过去除单词末尾的
s实现单复数统一匹配,若有更复杂的复数规则(如potatoes→potato),可扩展该逻辑 - 窗口检测:对每个匹配到的目标词,检查其前后
window_size个单词范围内是否存在否定词,不存在则计数 - 结果输出:返回每个目标词的有效计数
运行后会得到结果:
hello: 0 apple: 1
对应文本中只有I love apples, apple are my favorite fruit.里的apple/apples未被否定,其余两处都在don't/not的影响范围内,不计入统计。
内容的提问来源于stack exchange,提问作者Rollo99
相关产品推荐
相关产品推荐

