You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用WordNet进行NLP词形还原时触发[E1041]错误:输入为NoneType

问题分析与解决:文本预处理触发spaCy [E1041]错误

问题场景

定义了如下文本预处理函数pre_process:

import nltk
import unicodedata
from nltk.tokenize import TweetTokenizer
from nltk.corpus import wordnet
from nltk.stem import WordNetLemmatizer
from nltk.corpus import stopwords
from spellchecker import SpellChecker

# Tokenisation
tokenizer = TweetTokenizer() # Initialise an object of the TweetTokenizer class

# Lemmatization
nltk.download('wordnet') # Download lemmatizers from NLTK
lemmatizer = WordNetLemmatizer() # Initialise an object of the WordNetLemmatizer class 

# Stopwords 
nltk.download('stopwords') # Downloand stopwords from NLTK

# Spellchecker
spell_check = SpellChecker() # Initialise an object of the SpellChecker class

nlp = spacy.load("en_core_web_sm") # Load package "en_core_web_sm" from spacy

def pre_process(character_text):
    """Pre-process all the concatenated lines of a character, 
    using tokenization, spelling normalization and other techniques.

    ::character_text:: a string with all of one character's lines
    """

    initial_tokens = tokenizer.tokenize(character_text) # Tokenization with TweetTokenizer
    pos_tagged = nltk.pos_tag(initial_tokens) # Part of speech tagger to tag the list of tokens

    stop_words = set(stopwords.words("english")) # Initialise a set that stores common english     stopwords

    punc_cat = set(["Pc", "Pd", "Ps", "Pe", "Pi", "Pf", "Po"]) # Check punctuation in tokens
    # Taken from the NLTK's module for POS tagging using CRFSuite official documentation

    tokens = []
           
    for token, tag in pos_tagged:
        if token == "_EOL_": # If end of sentence 
             tokens.append(token)
            
        if unicodedata.category(token[0]) not in punc_cat: # If a token is not a punctuation letter
        # The unicodedata.category method returns the general category assigned to the character as a string
        
            misspelled = spell_check.unknown(token) # Check and correct mispelled words 
            if len(misspelled)!= 0:
                token = spell_check.correction(token)

            if token not in stop_words: # If a token is not a stopword
                # Check for the common word POS-taggers
                if tag[0] == 'V':
                    tag_type = wordnet.VERB
                elif tag[0] == 'J':
                    tag_type = wordnet.ADJ
                elif tag[0] == 'N':
                    tag_type = wordnet.NOUN
                elif tag[0] == 'R':
                    tag_type = wordnet.ADV
                else:
                    tag_type = None
            
                if tag_type is not None:  
                    token = lemmatizer.lemmatize(token, tag_type) # Lemmatization with Wordnet 
                
                tokens.append(token)
                
    return tokens

运行以下代码生成训练语料时:

# create list of pairs of (character name, pre-processed character) 
training_corpus = [(name, pre_process(doc)) for name, doc in sorted(train_character_docs.items())]
train_labels = [name for name, doc in training_corpus]

触发错误:[E1041] Expected a string, Doc, or bytes as input, but got: <class 'NoneType'>

已尝试通过if tag_type is not None:过滤词形还原的无效输入、删除数据集空行,但问题未解决。

错误根源

错误的核心原因是拼写校正函数可能返回None:当spell_check.correction(token)遇到完全无法识别的字符/字符串时,会返回None值。后续代码未对这个None值做过滤,直接将其添加到tokens列表中。如果后续流程(比如通过加载的spaCy nlp对象处理这些tokens)将None作为输入传递给spaCy的API,就会触发[E1041]错误。

另外,原函数中if token == "_EOL_":未使用elif或continue,导致即使是_EOL_也会进入后续的标点判断逻辑,属于冗余逻辑,但不是触发当前错误的直接原因。

解决方案

修改pre_process函数,在拼写校正后添加None值检查,过滤无效的None token;同时优化逻辑判断结构,避免冗余处理:

import nltk
import unicodedata
from nltk.tokenize import TweetTokenizer
from nltk.corpus import wordnet
from nltk.stem import WordNetLemmatizer
from nltk.corpus import stopwords
from spellchecker import SpellChecker
import spacy

# Tokenisation
tokenizer = TweetTokenizer()

# Lemmatization
nltk.download('wordnet')
lemmatizer = WordNetLemmatizer()

# Stopwords 
nltk.download('stopwords')

# Spellchecker
spell_check = SpellChecker()

nlp = spacy.load("en_core_web_sm")

def pre_process(character_text):
    """Pre-process all the concatenated lines of a character, 
    using tokenization, spelling normalization and other techniques.

    ::character_text:: a string with all of one character's lines
    """
    # 提前处理空输入
    if not character_text or not character_text.strip():
        return []

    initial_tokens = tokenizer.tokenize(character_text)
    pos_tagged = nltk.pos_tag(initial_tokens)

    stop_words = set(stopwords.words("english"))
    punc_cat = set(["Pc", "Pd", "Ps", "Pe", "Pi", "Pf", "Po"])

    tokens = []
           
    for token, tag in pos_tagged:
        if token == "_EOL_":
             tokens.append(token)
             # 跳过后续判断,避免冗余逻辑
             continue
            
        if unicodedata.category(token[0]) not in punc_cat:
            # 修正spell_check.unknown的参数:该方法接受可迭代对象,原代码传单个token会被拆解为字符
            misspelled = spell_check.unknown([token])
            if misspelled:
                corrected_token = spell_check.correction(token)
                # 检查校正结果,非空才替换,否则直接跳过该token
                if corrected_token is not None:
                    token = corrected_token
                else:
                    continue

            if token not in stop_words:
                # 映射POS标签到wordnet类型
                tag_type = None
                if tag[0] == 'V':
                    tag_type = wordnet.VERB
                elif tag[0] == 'J':
                    tag_type = wordnet.ADJ
                elif tag[0] == 'N':
                    tag_type = wordnet.NOUN
                elif tag[0] == 'R':
                    tag_type = wordnet.ADV
                
                if tag_type is not None:  
                    token = lemmatizer.lemmatize(token, tag_type)
                
                tokens.append(token)
                
    return tokens

关键修改点:

  1. 增加输入为空的判断,直接返回空列表,避免后续处理空字符串产生异常
  2. 修正spell_check.unknown的调用参数:该方法接受可迭代对象(如列表),原代码传入单个token会被拆解为字符,导致误判
  3. 拼写校正后添加None检查,无法校正的token直接跳过,不加入结果列表
  4. 将_EOL_的判断改为if+continue,避免冗余的标点判断逻辑

内容的提问来源于stack exchange,提问作者salkyna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 06:05:20