正则表达式小写文本时无法保留K.A.类首字母缩写的问题求助
问题分析
原代码中re.sub(r"\w+", ...)的正则表达式\w+仅匹配字母、数字和下划线组成的序列,而K.A.包含点号,会被拆分为K、A两个独立的匹配项。这两个单字符不在matches列表(列表中是完整的K.A.)里,因此被转成小写,最终出现k.a.的结果。
解决方案
方法1:调整匹配正则,覆盖带点缩写
修改替换时的正则表达式,使其能完整匹配带点的首字母缩写,同时结合集合加速匹配判断:
# -*- coding: utf-8 -*- #!/usr/bin/env python from __future__ import unicode_literals import re text = "This sentence contains ADS, NASA and K.A. as acronymns." # 匹配首字母缩写的正则 pattern = r'[A-Z][a-zA-Z]*[A-Z]|(?:[A-Z]\.)+' # 转成集合,提升查找效率 acronyms = set(re.findall(pattern, text)) def process_word(match): word = match.group() return word if word in acronyms else word.lower() # 修改替换正则:优先匹配带点缩写,再匹配普通单词 text2 = re.sub(r'(?:[A-Z]\.)+|\w+', process_word, text) print(text) print(text2) print(acronyms)
输出结果:
This sentence contains ADS, NASA and K.A. as acronymns. this sentence contains ADS, NASA and K.A. as acronymns. {'ADS', 'NASA', 'K.A.'}
方法2:占位符替换法
先将所有缩写替换为临时占位符,转小写后再还原原缩写:
# -*- coding: utf-8 -*- #!/usr/bin/env python from __future__ import unicode_literals import re text = "This sentence contains ADS, NASA and K.A. as acronymns." pattern = r'[A-Z][a-zA-Z]*[A-Z]|(?:[A-Z]\.)+' acronyms = re.findall(pattern, text) # 建立占位符与原缩写的映射 placeholder_map = {} for idx, acro in enumerate(acronyms): placeholder = f"__ACRO_{idx}__" placeholder_map[placeholder] = acro text = text.replace(acro, placeholder) # 整体转小写 text_lower = text.lower() # 还原原缩写 for placeholder, original in placeholder_map.items(): text_lower = text_lower.replace(placeholder.lower(), original) print(text_lower)
输出结果:
this sentence contains ADS, NASA and K.A. as acronymns.
内容的提问来源于stack exchange,提问作者Programmer_nltk
相关产品推荐
相关产品推荐

