调用.lower()方法处理标注文本时出现NoneType错误的排查与解决
问题描述
我有一个包含488篇标注文章的列表,想要对词元(lemma)应用.lower()方法,却遇到错误:AttributeError: 'NoneType' object has no attribute 'lower'。
相关代码如下:
file = open("Guardian_Syria_text.csv", mode="r", encoding='utf-8-sig') data = list(csv.reader(file, delimiter=",")) file.close pickle.dump(data, open('List.p', 'wb')) stanza.download('en') nlp = stanza.Pipeline(lang='en', processors='tokenize,lemma,POS', use_gpu=True) data_list = pickle.load(open('List.p', 'rb')) new_list = [] for article in data_list: a = nlp(str(article)) new_list.append(a) pickle.dump(new_list, open('Annotated.p', 'wb')) annot_data = pickle.load(open('Annotated.p', 'rb')) pos_tags = {'NOUN', 'VERB', 'ADJ', 'ADV', 'X'} lemmas = [] for article in annot_data: art_tokens = [w.text for s in article.sentences for w in s.words] art_lemmas = [w.lemma.lower() for s in article.sentences for w in s.words if w.upos in pos_tags] lemmas.append(art_lemmas)
我检查了变量annot_data是否为None(执行print(annot_data is None)),返回结果为False。我尝试通过clean = [x for x in annot_data if x != None]清理变量,但清理后的clean长度仍为488,使用该变量执行代码仍出现相同错误。请问这个NoneType对象来自哪里,该如何避免?
问题原因与解决办法
原因
报错的根源不是annot_data本身为None,而是部分词的w.lemma字段是None。Stanza在处理特殊文本(比如无意义符号、格式异常的词汇)时,可能无法生成有效词元,导致w.lemma返回None,此时调用.lower()就会触发AttributeError。
解决办法
修改生成词元列表的代码,提前处理w.lemma为None的情况,两种方案可选:
方案1:过滤掉lemma为None的词
art_lemmas = [w.lemma.lower() for s in article.sentences for w in s.words if w.upos in pos_tags and w.lemma is not None]
方案2:给None的lemma设置默认值(比如空字符串)
art_lemmas = [(w.lemma.lower() if w.lemma is not None else '') for s in article.sentences for w in s.words if w.upos in pos_tags]
额外优化建议
- 若
data中每一行本身就是单篇文章文本,无需将article转成字符串再传入Stanza,直接用nlp(article)即可(注意处理空文本),转字符串可能引入多余符号(如列表方括号、引号),增加处理负担。 - 文件操作推荐用
with语句,自动关闭文件避免资源泄漏:
with open("Guardian_Syria_text.csv", mode="r", encoding='utf-8-sig') as file: data = list(csv.reader(file, delimiter=","))
内容的提问来源于stack exchange,提问作者Linda Brck
相关产品推荐
相关产品推荐

