Python正则反向引用数字引用异常,无法提取目标子串问题
正则反向引用解惑与重复地名去重方案
问题梳理
用difflib做匹配时碰到两个糟心问题:
- 明明匹配单个单词,结果返回带重复词的双词地名(比如示例里的
_Days_Bay Bay) - 尝试用正则反向引用处理重复时,搞不清
\0/\1/\2的输出逻辑,想要提取Bay some of the time却没摸准门道,还疑惑r'\0'为啥会有那样的输出
你的测试代码
# removeDupeWords.py --- test to remove double words eg "The sun shines in,_Days_Bay Bay some of the time" import re testString = "The sun shines in,_Days_Bay Bay some of the time" # regex to capture comma to space of testString e.g ',_Days_Bay' refRegex = '(,\S+)' # regex to capture everything after e.g 'Bay some of the time' afterRegex = '(,\S+)(.*)' refString = re.search(refRegex, testString).group(0) # print(refString) afterString = re.sub(afterRegex, r'\2', testString) print(afterString)
你观察到的反向引用输出
The sun shines in The sun shines in,_Days_Bay The sun shines in Bay some of the time
问题解析与解决办法
1. 先搞懂\0到底是什么
你疑惑\0的输出,其实在Python正则的替换字符串里:
\1/\2对应第1、2个捕获组的内容\0(或者\g<0>)代表整个被匹配到的子串,不是空字符也不存在“分组0”的说法
你的afterRegex = '(,\S+)(.*)'会匹配,_Days_Bay Bay some of the time这一段。如果用r'\0'替换,等于把这段替换成它自己,结果应该和原字符串一致,但你得到的是The sun shines in——这大概率是测试时把r'\0'和空字符串搞混了,或者没加r前缀导致\0被解析成了空转义字符。
2. 直接拿到你想要的 Bay some of the time
如果只是想提取这部分内容,不用re.sub,直接用re.search抓第二个捕获组就行:
import re testString = "The sun shines in,_Days_Bay Bay some of the time" afterRegex = '(,\S+)(.*)' match_result = re.search(afterRegex, testString) if match_result: print(match_result.group(2)) # 输出: Bay some of the time
3. 解决重复单词的核心需求
你本来是要处理_Days_Bay Bay这种重复地名,直接针对重复词写正则更高效:
import re testString = "The sun shines in,_Days_Bay Bay some of the time" # 匹配“单词 + 空格 + 相同单词”的结构 dupe_pattern = r'(\b\w+)\s+\1\b' # 替换成单个单词 cleaned_string = re.sub(dupe_pattern, r'\1', testString) print(cleaned_string) # 输出:The sun shines in,_Days_Bay some of the time
如果是要提取重复词之后的内容,调整正则即可:
dupe_pattern = r'(\b\w+)\s+\1\b(.*)' match_result = re.search(dupe_pattern, testString) if match_result: print(match_result.group(2)) # 输出: some of the time
内容的提问来源于stack exchange,提问作者Dave
相关产品推荐
相关产品推荐

