Python实现否定词与后续单词拼接的方法,例将not apple转为not_apple
Python否定词与后续单词拼接实现方案
方法1:正则表达式实现(推荐,代码简洁高效)
直接通过正则匹配否定词+空格+后续单词的结构,批量替换为下划线拼接的格式,适合绝大多数场景。
import re # 自定义否定词集合,可按需增减 NEG_WORDS = {"no", "not", "never", "none", "nobody", "nowhere", "neither", "nor"} # 构建正则匹配规则,re.IGNORECASE表示忽略大小写匹配 pattern = re.compile(r"\b(" + "|".join(NEG_WORDS) + r")\s+(\w+)\b", re.IGNORECASE) # 测试示例 raw_text = "This is not apple" processed_text = pattern.sub(r"\1_\2", raw_text) print(processed_text) # 输出:This is not_apple
方法2:分词遍历实现(适合需要自定义复杂逻辑的场景)
通过拆分句子为单词列表遍历处理,灵活度更高,可适配标点、特殊字符等复杂文本场景。
NEG_WORDS = {"no", "not", "never", "none", "nobody", "nowhere", "neither", "nor"} def concat_neg_with_next_word(text: str) -> str: words = text.split() res = [] idx = 0 word_count = len(words) while idx < word_count: # 检测到否定词且存在后续单词时拼接 if words[idx].lower() in NEG_WORDS and idx < word_count - 1: res.append(f"{words[idx]}_{words[idx+1]}") idx += 2 else: res.append(words[idx]) idx += 1 return " ".join(res) # 测试示例 raw_text = "This is not apple" processed_text = concat_neg_with_next_word(raw_text) print(processed_text) # 输出:This is not_apple
可选调整项
- 不需要忽略大小写匹配时,删掉正则参数
re.IGNORECASE,或遍历匹配时去掉.lower()即可 - 可直接修改
NEG_WORDS集合增删需要识别的否定词 - 若需要处理带标点的文本,可先通过字符串替换或专业分词库预处理后再执行上述逻辑
内容的提问来源于stack exchange,提问作者Staw6874
相关产品推荐
相关产品推荐

