使用Python正则与termcolor高亮字符串点号致Jupyter崩溃求助
正则匹配点号(.)导致Jupyter内核崩溃问题排查与解决
问题描述
我编写了一段代码,通过正则表达式识别文本模式,借助termcolor库将匹配内容替换为高亮版本。多数模式下代码运行正常,但当匹配点号(.)时,Jupyter内核会无限运行后崩溃。
代码示例
import re from termcolor import colored,cprint text = "This is a sample text......" pattern = re.compile(r"\.") patternlist = pattern.findall(text) # print(patternlist) replacelist = [colored(i,"black", "on_yellow", attrs=["bold"]) for i in patternlist] print(replacelist) patterns = [i for i in zip(patternlist,replacelist)] print(patterns) for pattern, replacement in patterns: text = re.sub(pattern, replacement, text) print(text)
问题细节
使用pattern = re.compile(r"\.")时,findall能正常返回预期结果['.', '.', '.', '.', '.', '.'],预期得到高亮后的文本,但实际无输出,Jupyter内核无限运行后崩溃。
解决方案
问题根源
核心问题出在循环替换的逻辑上:
- 你用
findall拿到6个点号,生成了6组(原字符,高亮字符)的替换对 - 循环中执行
re.sub(pattern, replacement, text)时,这里的pattern是字符串.,而非之前编译的正则表达式 - 在正则语法中,
.是匹配任意字符的通配符,不是匹配点号本身。第一次替换后,文本中插入了termcolor的ANSI控制字符(如\x1b[30;1;43m.\x1b[0m),后续循环里re.sub(".", replacement, text)会匹配这些控制字符的每一个字符,不断重复替换导致文本无限膨胀,最终内核崩溃。
修复方式
方式一:复用编译好的正则一次性替换(推荐)
直接用编译完成的正则表达式,配合sub的回调函数一次性替换所有匹配项,逻辑简洁高效:
import re from termcolor import colored,cprint text = "This is a sample text......" pattern = re.compile(r"\.") # 用lambda回调处理每个匹配项的高亮 highlighted_text = pattern.sub(lambda match: colored(match.group(), "black", "on_yellow", attrs=["bold"]), text) print(highlighted_text)
方式二:对字符串模式转义并限制替换次数
如果需要保留循环逻辑,需用re.escape()将字符串.转义为正则字面量,同时加上count=1避免重复替换同一位置:
import re from termcolor import colored,cprint text = "This is a sample text......" pattern = re.compile(r"\.") patternlist = pattern.findall(text) replacelist = [colored(i,"black", "on_yellow", attrs=["bold"]) for i in patternlist] patterns = zip(patternlist, replacelist) for pattern_str, replacement in patterns: # 转义模式字符串,确保匹配字面量点号,且仅替换一次 text = re.sub(re.escape(pattern_str), replacement, text, count=1) print(text)
内容的提问来源于stack exchange,提问作者Nithin Mohan
相关产品推荐
相关产品推荐

