You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行聊天机器人代码时触发TypeError错误求助

问题:聊天机器人运行触发TypeError: int()参数不能为NoneType

开发聊天机器人时,运行代码持续触发如下错误:
TypeError: int() argument must be a string, a bytes-like object or a real number, not 'NoneType'

代码片段

def pre_process_data(data):
    tokens = nltk.word_tokenize(data)
    tokens = [word.lower() for word in tokens]
    stop_words = set(stopwords.words('english'))
    tokens = [
        word for word in tokens if word not in stop_words and word not in string.punctuation]
    lemmatizer = WordNetLemmatizer()
    tokens = [lemmatizer.lemmatize(word) for word in tokens]
    return tokens


def predict_answer(model, tokenizer, question):
    question = pre_process_data(question)
    sequence = tokenizer.texts_to_sequences([question])
    padded_sequence = pad_sequences(
        sequence, maxlen=sequence_length, padding='post', truncating='post')
    prediction = model.predict(padded_sequence)[0]
    index = numpy.argmax(prediction)
    answer = tokenizer.index_word[index]
    return answer


while True:
    question = input('You: ')
    answer = predict_answer(model, tokenizer, question)
    print('Chatbot:', answer)

完整报错回溯

Traceback (most recent call last):
  File "D:\Chatbot\main.py", line 119, in <module>
    answer = predict_answer(model, tokenizer, question)
  File "D:\Chatbot\main.py", line 109, in predict_answer
    padded_sequence = pad_sequences(
  File "D:\Chatbot\.venv\lib\site-packages\keras\src\utils\sequence_utils.py", line 125, in pad_sequences
    trunc = np.asarray(trunc, dtype=dtype)
TypeError: int() argument must be a string, a bytes-like object or a real number, not 'NoneType'

问题原因

  1. 直接原因:sequence_length变量未被正确初始化,值为None。pad_sequences函数要求maxlen参数为整数,无法将None转换为整数类型,因此触发报错。
  2. 潜在风险:若输入内容经预处理后变为空列表(比如输入全是停用词或标点符号),tokenizer.texts_to_sequences([question])会生成全0序列,后续numpy.argmax(prediction)可能返回索引0,而tokenizer.index_word[0]通常不存在,会引发KeyError。

解决方案

1. 修复sequence_length变量

确保sequence_length是一个与训练模型时一致的整数,比如在代码开头定义:

sequence_length = 30  # 替换为你训练模型时使用的实际序列长度

2. 处理空输入场景

修改predict_answer函数,增加对预处理后空列表的判断,同时兼容索引不存在的情况:

def predict_answer(model, tokenizer, question):
    question = pre_process_data(question)
    if not question:  # 预处理后无有效词汇
        return "Sorry, I don't get that."
    sequence = tokenizer.texts_to_sequences([question])
    padded_sequence = pad_sequences(
        sequence, maxlen=sequence_length, padding='post', truncating='post')
    prediction = model.predict(padded_sequence)[0]
    index = numpy.argmax(prediction)
    # 使用get方法避免KeyError
    answer = tokenizer.index_word.get(index, "Sorry, I can't answer that right now.")
    return answer

3. 验证tokenizer有效性

确保tokenizer是基于训练数据拟合生成的,能够识别预处理后的词汇,避免生成无有效索引的序列。

内容的提问来源于stack exchange,提问作者Reina297

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 03:27:36