You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为文本特定词语添加引用编号:正则匹配边缘问题求助

问题描述

需要为文本中的特定词语添加(数字)格式的引用编号,现有Python正则代码仅能得到部分正确结果,遇到带形容词的相同词语或词语带后缀的情况时失效。

  • 错误输出:This is a first sample (3) (1) and this is a second sample (3) (2)
  • 期望输出:This is a first sample (1) and this is a second sample (2)

尝试的代码:

import re

text = "This is a first sample and this is a second sample."
words_to_number = {"first sample": 1, "second sample": 2, "sample": 3}

for keyword, number in words_to_number.items():
    pattern = r"\b"+keyword+r"\b"
    text = re.sub(pattern, keyword+" ("+str(number)+")", text)

print(text)
问题原因
  1. 字典遍历顺序不合理:原代码按插入顺序处理关键词,短关键词sample最后处理时,会匹配到已替换完成的长短语中的sample单词,导致重复添加编号。
  2. 正则匹配无优先级:短关键词的正则规则会匹配长短语的子部分,破坏已完成的替换结果。
解决方案

核心思路是优先匹配更长的关键词,避免短关键词干扰长短语的匹配。具体实现:

  • 对关键词按长度从长到短排序,确保长短语先被处理
  • 保留正则边界匹配,保证仅匹配完整的目标词语/短语

修正后的代码:

import re

text = "This is a first sample and this is a second sample."
words_to_number = {"first sample": 1, "second sample": 2, "sample": 3}

# 按关键词长度倒序排序,优先处理长短语
sorted_keywords = sorted(words_to_number.items(), key=lambda x: len(x[0]), reverse=True)

for keyword, number in sorted_keywords:
    # 转义关键词中的特殊正则字符,避免语法冲突
    pattern = re.compile(rf"\b{re.escape(keyword)}\b")
    text = pattern.sub(f"{keyword} ({number})", text)

print(text)
代码说明
  • sorted(..., key=lambda x: len(x[0]), reverse=True):将关键词按字符长度从长到短排序,确保first sample、second sample这类长短语先被替换,不会被后续的sample匹配干扰。
  • re.escape(keyword):转义关键词中的特殊正则字符(比如.、*等),避免正则语法错误。
  • re.compile预编译正则:提升替换效率,尤其适合处理大量文本的场景。

内容的提问来源于stack exchange,提问作者bloodymoonmate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 21:35:17