Python中列表元素Unicode规范化移除重音的报错解决方法
问题根因
报错TypeError: normalize() argument 2 must be str, not list来自几处代码错误:
unicodedata.normalize()仅支持字符串类型入参,原代码直接将分词后得到的单词列表words_w_stopwords作为参数传入,类型不匹配触发报错- 原代码导入时写的是
from unicodedata import normalize,调用时却使用unicodedata.normalize写法,会额外触发名称引用错误;同时去重音后直接将结果拼接为长字符串再逐字符遍历,会把所有单词拆分为零散字母,完全破坏文本的词结构。另外自定义干扰词列表parasites未提前定义,运行时也会报错。
修复后完整代码
import string import unicodedata import nltk from french_lefff_lemmatizer.french_lefff_lemmatizer import FrenchLefffLemmatizer # 提前下载nltk依赖 nltk.download('wordnet') nltk.download('punkt') nltk.download('stopwords') # 加载法语停用词 french_stopwords = nltk.corpus.stopwords.words('french') # 初始化词形还原器 lemmatizer = FrenchLefffLemmatizer() # 自定义需要过滤的干扰词列表,根据实际业务需求填充内容 parasites = [] def preprocessing(affaires): preprocess_list = [] for sentence in affaires : # 转小写、移除标点符号 sentence_w_punct = "".join([i.lower() for i in sentence if i not in string.punctuation]) # 分词 tokenize_sentence = nltk.tokenize.word_tokenize(sentence_w_punct) # 过滤通用法语停用词 words_w_stopwords = [i for i in tokenize_sentence if i not in french_stopwords] # 逐单词移除重音:遍历列表,对每个字符串类型的单词单独执行Unicode规范化 no_accent = [] for word in words_w_stopwords: word_no_accent = ''.join( c for c in unicodedata.normalize('NFD', word) if unicodedata.category(c) != 'Mn' ) no_accent.append(word_no_accent) # 过滤自定义干扰词 remove_parasites = [j for j in no_accent if j not in parasites] # 词形还原 words_lemmatize = [lemmatizer.lemmatize(w) for w in remove_parasites] # 拼接为清洗后的完整句子 sentence_clean = ' '.join(words_lemmatize) preprocess_list.append(sentence_clean) return preprocess_list # 执行预处理 df["nom_affaire_clean"] = preprocessing(df["nom_affaire"]) # 调整列顺序 cln = df.pop("nom_affaire_clean") df.insert(1, 'nom_affaire_clean', cln )
核心修改说明
- 调整重音移除逻辑的作用粒度:不再将整个单词列表传入
normalize函数,改为遍历列表,对每个单词(字符串类型)单独做规范化处理,从根源解决类型不匹配问题 - 保留词结构:去重音处理后保留单词列表形态,不会把单词拆成单个字符
- 精简冗余导入:删除重复的停用词加载代码,统一
unicodedata的导入和调用写法 - 补全缺失变量:提前声明自定义干扰词列表
parasites,避免未定义错误
内容的提问来源于stack exchange,提问作者Julie-Anne
相关产品推荐
相关产品推荐

