You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python字符串列表归一化:去除冗余标点实现格式统一

解决方案

一、基础版:移除所有指定标点并转小写

如果只需要移除指定标点(无论位置)并统一为小写,使用正则表达式是最高效的方式,适合处理大规模列表:

import re

def word_normalizer(word):
    # 定义需要移除的标点字符集(包含你列出的所有标点,加上例子中的问号)
    punctuation_pattern = r"['\";:,.&()?]"
    # 替换所有匹配的标点为空字符串,再转小写
    cleaned_word = re.sub(punctuation_pattern, "", word)
    return cleaned_word.lower()

二、进阶版:额外处理所有格与规则复数

如果需要将Cognition's、Cognitions这类形式归一化为Cognition,可以在移除标点后添加规则复数/所有格的处理逻辑:

import re

def word_normalizer(word):
    # 1. 先转小写统一格式
    word_lower = word.lower()
    # 2. 移除所有指定标点
    punctuation_pattern = r"['\";:,.&()?]"
    cleaned_word = re.sub(punctuation_pattern, "", word_lower)
    # 3. 处理规则复数/所有格:去掉末尾的s(排除单字母"s"的情况)
    if len(cleaned_word) > 1 and cleaned_word.endswith('s'):
        cleaned_word = cleaned_word[:-1]
    return cleaned_word

测试效果

用你的示例列表测试进阶版函数:

word_list = ["Alzheimer", "Alzheimer's", "Alzheimer.", "Alzheimer?","Cognition.", "Cognition's", "Cognitions", "Cognition"]
normalized_list = [word_normalizer(word) for word in word_list]
print(normalized_list)
# 输出:['alzheimer', 'alzheimer', 'alzheimer', 'alzheimer', 'cognition', 'cognition', 'cognition', 'cognition']

原函数的问题分析

你的原始函数无法正常工作,主要有三个问题:

  1. 循环提前终止:return语句写在for循环内部,导致循环只执行一次就返回结果,无法处理多个标点。
  2. strip的局限性:strip(punc)只能移除字符串首尾的标点,无法处理中间的标点(比如Alzheimer's中的单引号)。
  3. 空值风险:如果单词中没有第一个标点,new_word会保持为空字符串,最终返回空值,不符合预期。

内容的提问来源于stack exchange,提问作者Yixing Wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 16:01:11