使用Scikit-learn实现简易对话机器人时遭遇AttributeError问题求助
嘿,我来帮你排查这个问题!你遇到的AttributeError: 'list' object has no attribute 'lower'其实是因为给TfidfVectorizer传的数据格式不对,咱们一步步理清楚:
错误原因分析
你的preprocess_text函数已经能把单个字符串转换成分词后的列表,但你在generate_response里犯了一个小错误:你提前把用户输入和对话库的所有内容都用preprocess_text处理了一遍,得到的conversation_tokens是一个列表套列表的结构。而TfidfVectorizer.fit_transform()需要的输入是字符串组成的列表,它会自动调用你指定的tokenizer(也就是preprocess_text)来处理每个字符串。
现在你提前把文本转成了列表,Vectorizer在处理时会把每个子列表当成要处理的“文本”,试图调用字符串的lower()方法,可列表根本没有这个属性,自然就报错了。
修正后的代码方案
只需要修改generate_response函数,去掉提前预处理的步骤,直接给Vectorizer传原始字符串就行,它会帮你调用preprocess_text分词:
def generate_response(user_input, conversations): response = '' # 直接收集原始字符串:用户输入 + 对话库内容 conversation_texts = [user_input] + conversations # TfidfVectorizer会自动用preprocess_text处理每个字符串 tfidf_vectorizer = TfidfVectorizer(tokenizer=preprocess_text) tfidf_matrix = tfidf_vectorizer.fit_transform(conversation_texts) similarity_scores = cosine_similarity(tfidf_matrix[-1], tfidf_matrix) idx = similarity_scores.argsort()[0][-2] flat = similarity_scores.flatten() flat.sort() score = flat[-2] if score == 0: response = 'I apologize, I do not understand.' else: response = conversations[idx] return response
为什么这样改就对了?
TfidfVectorizer的tokenizer参数设计就是接收单个字符串,返回该字符串的分词列表。现在我们直接传原始字符串,Vectorizer会逐个把字符串传给preprocess_text处理,完全符合它的预期,就不会再出现类型错误了。
另外,你的preprocess_text里的类型检查做得很棒,能提前防止非字符串输入的问题,这个可以保留~
备注:内容来源于stack exchange,提问作者Juan Sparda

