运行聊天机器人代码时触发TypeError错误求助
问题:聊天机器人运行触发TypeError: int()参数不能为NoneType
开发聊天机器人时,运行代码持续触发如下错误:TypeError: int() argument must be a string, a bytes-like object or a real number, not 'NoneType'
代码片段
def pre_process_data(data): tokens = nltk.word_tokenize(data) tokens = [word.lower() for word in tokens] stop_words = set(stopwords.words('english')) tokens = [ word for word in tokens if word not in stop_words and word not in string.punctuation] lemmatizer = WordNetLemmatizer() tokens = [lemmatizer.lemmatize(word) for word in tokens] return tokens def predict_answer(model, tokenizer, question): question = pre_process_data(question) sequence = tokenizer.texts_to_sequences([question]) padded_sequence = pad_sequences( sequence, maxlen=sequence_length, padding='post', truncating='post') prediction = model.predict(padded_sequence)[0] index = numpy.argmax(prediction) answer = tokenizer.index_word[index] return answer while True: question = input('You: ') answer = predict_answer(model, tokenizer, question) print('Chatbot:', answer)
完整报错回溯
Traceback (most recent call last): File "D:\Chatbot\main.py", line 119, in <module> answer = predict_answer(model, tokenizer, question) File "D:\Chatbot\main.py", line 109, in predict_answer padded_sequence = pad_sequences( File "D:\Chatbot\.venv\lib\site-packages\keras\src\utils\sequence_utils.py", line 125, in pad_sequences trunc = np.asarray(trunc, dtype=dtype) TypeError: int() argument must be a string, a bytes-like object or a real number, not 'NoneType'
问题原因
- 直接原因:
sequence_length变量未被正确初始化,值为None。pad_sequences函数要求maxlen参数为整数,无法将None转换为整数类型,因此触发报错。 - 潜在风险:若输入内容经预处理后变为空列表(比如输入全是停用词或标点符号),
tokenizer.texts_to_sequences([question])会生成全0序列,后续numpy.argmax(prediction)可能返回索引0,而tokenizer.index_word[0]通常不存在,会引发KeyError。
解决方案
1. 修复sequence_length变量
确保sequence_length是一个与训练模型时一致的整数,比如在代码开头定义:
sequence_length = 30 # 替换为你训练模型时使用的实际序列长度
2. 处理空输入场景
修改predict_answer函数,增加对预处理后空列表的判断,同时兼容索引不存在的情况:
def predict_answer(model, tokenizer, question): question = pre_process_data(question) if not question: # 预处理后无有效词汇 return "Sorry, I don't get that." sequence = tokenizer.texts_to_sequences([question]) padded_sequence = pad_sequences( sequence, maxlen=sequence_length, padding='post', truncating='post') prediction = model.predict(padded_sequence)[0] index = numpy.argmax(prediction) # 使用get方法避免KeyError answer = tokenizer.index_word.get(index, "Sorry, I can't answer that right now.") return answer
3. 验证tokenizer有效性
确保tokenizer是基于训练数据拟合生成的,能够识别预处理后的词汇,避免生成无有效索引的序列。
内容的提问来源于stack exchange,提问作者Reina297
相关产品推荐
相关产品推荐

