使用WordNet进行NLP词形还原时触发[E1041]错误:输入为NoneType
问题分析与解决:文本预处理触发spaCy [E1041]错误
问题场景
定义了如下文本预处理函数pre_process:
import nltk import unicodedata from nltk.tokenize import TweetTokenizer from nltk.corpus import wordnet from nltk.stem import WordNetLemmatizer from nltk.corpus import stopwords from spellchecker import SpellChecker # Tokenisation tokenizer = TweetTokenizer() # Initialise an object of the TweetTokenizer class # Lemmatization nltk.download('wordnet') # Download lemmatizers from NLTK lemmatizer = WordNetLemmatizer() # Initialise an object of the WordNetLemmatizer class # Stopwords nltk.download('stopwords') # Downloand stopwords from NLTK # Spellchecker spell_check = SpellChecker() # Initialise an object of the SpellChecker class nlp = spacy.load("en_core_web_sm") # Load package "en_core_web_sm" from spacy def pre_process(character_text): """Pre-process all the concatenated lines of a character, using tokenization, spelling normalization and other techniques. ::character_text:: a string with all of one character's lines """ initial_tokens = tokenizer.tokenize(character_text) # Tokenization with TweetTokenizer pos_tagged = nltk.pos_tag(initial_tokens) # Part of speech tagger to tag the list of tokens stop_words = set(stopwords.words("english")) # Initialise a set that stores common english stopwords punc_cat = set(["Pc", "Pd", "Ps", "Pe", "Pi", "Pf", "Po"]) # Check punctuation in tokens # Taken from the NLTK's module for POS tagging using CRFSuite official documentation tokens = [] for token, tag in pos_tagged: if token == "_EOL_": # If end of sentence tokens.append(token) if unicodedata.category(token[0]) not in punc_cat: # If a token is not a punctuation letter # The unicodedata.category method returns the general category assigned to the character as a string misspelled = spell_check.unknown(token) # Check and correct mispelled words if len(misspelled)!= 0: token = spell_check.correction(token) if token not in stop_words: # If a token is not a stopword # Check for the common word POS-taggers if tag[0] == 'V': tag_type = wordnet.VERB elif tag[0] == 'J': tag_type = wordnet.ADJ elif tag[0] == 'N': tag_type = wordnet.NOUN elif tag[0] == 'R': tag_type = wordnet.ADV else: tag_type = None if tag_type is not None: token = lemmatizer.lemmatize(token, tag_type) # Lemmatization with Wordnet tokens.append(token) return tokens
运行以下代码生成训练语料时:
# create list of pairs of (character name, pre-processed character) training_corpus = [(name, pre_process(doc)) for name, doc in sorted(train_character_docs.items())] train_labels = [name for name, doc in training_corpus]
触发错误:[E1041] Expected a string, Doc, or bytes as input, but got: <class 'NoneType'>
已尝试通过if tag_type is not None:过滤词形还原的无效输入、删除数据集空行,但问题未解决。
错误根源
错误的核心原因是拼写校正函数可能返回None:当spell_check.correction(token)遇到完全无法识别的字符/字符串时,会返回None值。后续代码未对这个None值做过滤,直接将其添加到tokens列表中。如果后续流程(比如通过加载的spaCy nlp对象处理这些tokens)将None作为输入传递给spaCy的API,就会触发[E1041]错误。
另外,原函数中if token == "_EOL_":未使用elif或continue,导致即使是_EOL_也会进入后续的标点判断逻辑,属于冗余逻辑,但不是触发当前错误的直接原因。
解决方案
修改pre_process函数,在拼写校正后添加None值检查,过滤无效的None token;同时优化逻辑判断结构,避免冗余处理:
import nltk import unicodedata from nltk.tokenize import TweetTokenizer from nltk.corpus import wordnet from nltk.stem import WordNetLemmatizer from nltk.corpus import stopwords from spellchecker import SpellChecker import spacy # Tokenisation tokenizer = TweetTokenizer() # Lemmatization nltk.download('wordnet') lemmatizer = WordNetLemmatizer() # Stopwords nltk.download('stopwords') # Spellchecker spell_check = SpellChecker() nlp = spacy.load("en_core_web_sm") def pre_process(character_text): """Pre-process all the concatenated lines of a character, using tokenization, spelling normalization and other techniques. ::character_text:: a string with all of one character's lines """ # 提前处理空输入 if not character_text or not character_text.strip(): return [] initial_tokens = tokenizer.tokenize(character_text) pos_tagged = nltk.pos_tag(initial_tokens) stop_words = set(stopwords.words("english")) punc_cat = set(["Pc", "Pd", "Ps", "Pe", "Pi", "Pf", "Po"]) tokens = [] for token, tag in pos_tagged: if token == "_EOL_": tokens.append(token) # 跳过后续判断,避免冗余逻辑 continue if unicodedata.category(token[0]) not in punc_cat: # 修正spell_check.unknown的参数:该方法接受可迭代对象,原代码传单个token会被拆解为字符 misspelled = spell_check.unknown([token]) if misspelled: corrected_token = spell_check.correction(token) # 检查校正结果,非空才替换,否则直接跳过该token if corrected_token is not None: token = corrected_token else: continue if token not in stop_words: # 映射POS标签到wordnet类型 tag_type = None if tag[0] == 'V': tag_type = wordnet.VERB elif tag[0] == 'J': tag_type = wordnet.ADJ elif tag[0] == 'N': tag_type = wordnet.NOUN elif tag[0] == 'R': tag_type = wordnet.ADV if tag_type is not None: token = lemmatizer.lemmatize(token, tag_type) tokens.append(token) return tokens
关键修改点:
- 增加输入为空的判断,直接返回空列表,避免后续处理空字符串产生异常
- 修正
spell_check.unknown的调用参数:该方法接受可迭代对象(如列表),原代码传入单个token会被拆解为字符,导致误判 - 拼写校正后添加
None检查,无法校正的token直接跳过,不加入结果列表 - 将
_EOL_的判断改为if+continue,避免冗余的标点判断逻辑
内容的提问来源于stack exchange,提问作者salkyna
相关产品推荐
相关产品推荐

