You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scikit-learn实现简易对话机器人时遭遇AttributeError问题求助

Scikit-learn实现简易对话机器人时遭遇AttributeError问题求助

嘿,我来帮你排查这个问题!你遇到的AttributeError: 'list' object has no attribute 'lower'其实是因为给TfidfVectorizer传的数据格式不对,咱们一步步理清楚:

错误原因分析

你的preprocess_text函数已经能把单个字符串转换成分词后的列表,但你在generate_response里犯了一个小错误:你提前把用户输入和对话库的所有内容都用preprocess_text处理了一遍,得到的conversation_tokens是一个列表套列表的结构。而TfidfVectorizer.fit_transform()需要的输入是字符串组成的列表,它会自动调用你指定的tokenizer(也就是preprocess_text)来处理每个字符串。

现在你提前把文本转成了列表,Vectorizer在处理时会把每个子列表当成要处理的“文本”,试图调用字符串的lower()方法,可列表根本没有这个属性,自然就报错了。

修正后的代码方案

只需要修改generate_response函数,去掉提前预处理的步骤,直接给Vectorizer传原始字符串就行,它会帮你调用preprocess_text分词:

def generate_response(user_input, conversations):
    response = ''
    # 直接收集原始字符串:用户输入 + 对话库内容
    conversation_texts = [user_input] + conversations
    
    # TfidfVectorizer会自动用preprocess_text处理每个字符串
    tfidf_vectorizer = TfidfVectorizer(tokenizer=preprocess_text)
    tfidf_matrix = tfidf_vectorizer.fit_transform(conversation_texts)
    
    similarity_scores = cosine_similarity(tfidf_matrix[-1], tfidf_matrix)
    idx = similarity_scores.argsort()[0][-2]
    flat = similarity_scores.flatten()
    flat.sort()
    score = flat[-2]
    
    if score == 0:
        response = 'I apologize, I do not understand.'
    else:
        response = conversations[idx]
    return response

为什么这样改就对了?

TfidfVectorizer的tokenizer参数设计就是接收单个字符串,返回该字符串的分词列表。现在我们直接传原始字符串,Vectorizer会逐个把字符串传给preprocess_text处理,完全符合它的预期,就不会再出现类型错误了。

另外,你的preprocess_text里的类型检查做得很棒,能提前防止非字符串输入的问题,这个可以保留~

备注:内容来源于stack exchange,提问作者Juan Sparda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.23 10:45:29