You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中列表元素Unicode规范化移除重音的报错解决方法

问题根因

报错TypeError: normalize() argument 2 must be str, not list来自几处代码错误:

  1. unicodedata.normalize() 仅支持字符串类型入参,原代码直接将分词后得到的单词列表words_w_stopwords作为参数传入,类型不匹配触发报错
  2. 原代码导入时写的是from unicodedata import normalize,调用时却使用unicodedata.normalize写法,会额外触发名称引用错误;同时去重音后直接将结果拼接为长字符串再逐字符遍历,会把所有单词拆分为零散字母,完全破坏文本的词结构。另外自定义干扰词列表parasites未提前定义,运行时也会报错。
修复后完整代码
import string
import unicodedata
import nltk
from french_lefff_lemmatizer.french_lefff_lemmatizer import FrenchLefffLemmatizer

# 提前下载nltk依赖
nltk.download('wordnet')
nltk.download('punkt')
nltk.download('stopwords')
# 加载法语停用词
french_stopwords = nltk.corpus.stopwords.words('french')
# 初始化词形还原器
lemmatizer = FrenchLefffLemmatizer()
# 自定义需要过滤的干扰词列表,根据实际业务需求填充内容
parasites = []

def preprocessing(affaires):
    preprocess_list = []
    for sentence in affaires :
        # 转小写、移除标点符号
        sentence_w_punct = "".join([i.lower() for i in sentence if i not in string.punctuation])
        # 分词
        tokenize_sentence = nltk.tokenize.word_tokenize(sentence_w_punct)
        # 过滤通用法语停用词
        words_w_stopwords = [i for i in tokenize_sentence if i not in french_stopwords]
        # 逐单词移除重音:遍历列表,对每个字符串类型的单词单独执行Unicode规范化
        no_accent = []
        for word in words_w_stopwords:
            word_no_accent = ''.join(
                c for c in unicodedata.normalize('NFD', word)
                if unicodedata.category(c) != 'Mn'
            )
            no_accent.append(word_no_accent)
        # 过滤自定义干扰词
        remove_parasites = [j for j in no_accent if j not in parasites]
        # 词形还原
        words_lemmatize = [lemmatizer.lemmatize(w) for w in remove_parasites]
        # 拼接为清洗后的完整句子
        sentence_clean = ' '.join(words_lemmatize)
        preprocess_list.append(sentence_clean)

    return preprocess_list

# 执行预处理
df["nom_affaire_clean"] = preprocessing(df["nom_affaire"])
# 调整列顺序
cln = df.pop("nom_affaire_clean")
df.insert(1, 'nom_affaire_clean', cln )
核心修改说明
  • 调整重音移除逻辑的作用粒度:不再将整个单词列表传入normalize函数,改为遍历列表,对每个单词(字符串类型)单独做规范化处理,从根源解决类型不匹配问题
  • 保留词结构:去重音处理后保留单词列表形态,不会把单词拆成单个字符
  • 精简冗余导入:删除重复的停用词加载代码,统一unicodedata的导入和调用写法
  • 补全缺失变量:提前声明自定义干扰词列表parasites,避免未定义错误

内容的提问来源于stack exchange,提问作者Julie-Anne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 12:57:11