Python字符串列表归一化:去除冗余标点实现格式统一
解决方案
一、基础版:移除所有指定标点并转小写
如果只需要移除指定标点(无论位置)并统一为小写,使用正则表达式是最高效的方式,适合处理大规模列表:
import re def word_normalizer(word): # 定义需要移除的标点字符集(包含你列出的所有标点,加上例子中的问号) punctuation_pattern = r"['\";:,.&()?]" # 替换所有匹配的标点为空字符串,再转小写 cleaned_word = re.sub(punctuation_pattern, "", word) return cleaned_word.lower()
二、进阶版:额外处理所有格与规则复数
如果需要将Cognition's、Cognitions这类形式归一化为Cognition,可以在移除标点后添加规则复数/所有格的处理逻辑:
import re def word_normalizer(word): # 1. 先转小写统一格式 word_lower = word.lower() # 2. 移除所有指定标点 punctuation_pattern = r"['\";:,.&()?]" cleaned_word = re.sub(punctuation_pattern, "", word_lower) # 3. 处理规则复数/所有格:去掉末尾的s(排除单字母"s"的情况) if len(cleaned_word) > 1 and cleaned_word.endswith('s'): cleaned_word = cleaned_word[:-1] return cleaned_word
测试效果
用你的示例列表测试进阶版函数:
word_list = ["Alzheimer", "Alzheimer's", "Alzheimer.", "Alzheimer?","Cognition.", "Cognition's", "Cognitions", "Cognition"] normalized_list = [word_normalizer(word) for word in word_list] print(normalized_list) # 输出:['alzheimer', 'alzheimer', 'alzheimer', 'alzheimer', 'cognition', 'cognition', 'cognition', 'cognition']
原函数的问题分析
你的原始函数无法正常工作,主要有三个问题:
- 循环提前终止:
return语句写在for循环内部,导致循环只执行一次就返回结果,无法处理多个标点。 - strip的局限性:
strip(punc)只能移除字符串首尾的标点,无法处理中间的标点(比如Alzheimer's中的单引号)。 - 空值风险:如果单词中没有第一个标点,
new_word会保持为空字符串,最终返回空值,不符合预期。
内容的提问来源于stack exchange,提问作者Yixing Wang
相关产品推荐
相关产品推荐

