You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式小写文本时无法保留K.A.类首字母缩写的问题求助

问题分析

原代码中re.sub(r"\w+", ...)的正则表达式\w+仅匹配字母、数字和下划线组成的序列,而K.A.包含点号,会被拆分为K、A两个独立的匹配项。这两个单字符不在matches列表(列表中是完整的K.A.)里,因此被转成小写,最终出现k.a.的结果。

解决方案

方法1:调整匹配正则,覆盖带点缩写

修改替换时的正则表达式,使其能完整匹配带点的首字母缩写,同时结合集合加速匹配判断:

# -*- coding: utf-8 -*-
#!/usr/bin/env python
from __future__ import unicode_literals

import re

text = "This sentence contains ADS, NASA and K.A. as acronymns."

# 匹配首字母缩写的正则
pattern = r'[A-Z][a-zA-Z]*[A-Z]|(?:[A-Z]\.)+'
# 转成集合,提升查找效率
acronyms = set(re.findall(pattern, text))

def process_word(match):
    word = match.group()
    return word if word in acronyms else word.lower()

# 修改替换正则:优先匹配带点缩写,再匹配普通单词
text2 = re.sub(r'(?:[A-Z]\.)+|\w+', process_word, text)

print(text)
print(text2)
print(acronyms)

输出结果:

This sentence contains ADS, NASA and K.A. as acronymns.
this sentence contains ADS, NASA and K.A. as acronymns.
{'ADS', 'NASA', 'K.A.'}

方法2:占位符替换法

先将所有缩写替换为临时占位符,转小写后再还原原缩写:

# -*- coding: utf-8 -*-
#!/usr/bin/env python
from __future__ import unicode_literals

import re

text = "This sentence contains ADS, NASA and K.A. as acronymns."
pattern = r'[A-Z][a-zA-Z]*[A-Z]|(?:[A-Z]\.)+'
acronyms = re.findall(pattern, text)

# 建立占位符与原缩写的映射
placeholder_map = {}
for idx, acro in enumerate(acronyms):
    placeholder = f"__ACRO_{idx}__"
    placeholder_map[placeholder] = acro
    text = text.replace(acro, placeholder)

# 整体转小写
text_lower = text.lower()

# 还原原缩写
for placeholder, original in placeholder_map.items():
    text_lower = text_lower.replace(placeholder.lower(), original)

print(text_lower)

输出结果:

this sentence contains ADS, NASA and K.A. as acronymns.

内容的提问来源于stack exchange,提问作者Programmer_nltk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 18:32:17