如何用Python NLTK实现用户输入与Excel问题数据的相似度匹配?
解决Python聊天机器人的问题匹配需求
嘿,我来帮你搞定这个问题!你已经成功把CSV数据读进来了,接下来核心就是问题匹配——既要处理用户输入的小差异(比如有没有问号、大小写不同),又要准确找到对应的问答条目。下面分步骤给你讲清楚怎么做:
第一步:优化数据存储与预处理
首先,我们需要把读取到的数据做标准化处理,这样能消除用户输入的格式差异。比如把问题里的标点去掉、转成小写,这样不管用户输入“伦敦在哪里?”还是“伦敦在哪里”,处理后都是一样的内容,方便匹配。
修改后的读取代码如下:
import csv import string # 读取CSV数据,存储为包含完整信息的列表(包含预处理后的问题) qa_dataset = [] with open("dataset.csv", encoding="utf-8") as csvfile: reader = csv.DictReader(csvfile) for row in reader: # 预处理问题:移除所有标点、转成小写 processed_question = row['Question'].translate(str.maketrans('', '', string.punctuation)).lower() qa_dataset.append({ 'QuestionID': row['QuestionID'], 'OriginalQuestion': row['Question'], 'ProcessedQuestion': processed_question, 'Answer': row['Answer'], 'Document': row['Document'] })
第二步:实现精确匹配函数
接下来写一个匹配函数,接收用户输入后,用同样的规则预处理,然后遍历数据集找到完全匹配的条目,返回完整数据:
def find_matching_answer(user_input): # 预处理用户输入,和存储的问题保持一致规则 processed_input = user_input.translate(str.maketrans('', '', string.punctuation)).lower() # 遍历数据集查找匹配项 for item in qa_dataset: if item['ProcessedQuestion'] == processed_input: return item # 返回包含QuestionID、Answer、Document的完整条目 # 没有找到匹配时返回提示 return {"error": "没有找到对应的问题哦"}
测试一下效果
# 测试带问号的输入 user_query1 = "伦敦在哪里?" print(find_matching_answer(user_query1)) # 输出:{'QuestionID': 'Q1', 'OriginalQuestion': '伦敦在哪里?', 'ProcessedQuestion': '伦敦在哪里', 'Answer': '在英国', 'Document': 'Google'} # 测试不带问号的输入 user_query2 = "伦敦在哪里" print(find_matching_answer(user_query2)) # 输出和上面完全一致 # 测试另一个问题 user_query3 = "球场上有多少足球运动员?" print(find_matching_answer(user_query3)) # 输出:{'QuestionID': 'Q2', 'OriginalQuestion': '球场上有多少足球运动员?', 'ProcessedQuestion': '球场上有多少足球运动员', 'Answer': '22', 'Document': 'Google'}
第三步:扩展模糊匹配(可选)
如果需要处理用户表述不同但意思相近的问题(比如用户问“伦敦在哪个国家”,想匹配到“伦敦在哪里?”),可以用模糊匹配库fuzzywuzzy来实现:
- 先安装依赖:
pip install fuzzywuzzy python-Levenshtein
- 编写模糊匹配函数:
from fuzzywuzzy import fuzz def find_fuzzy_matching_answer(user_input, threshold=80): processed_input = user_input.translate(str.maketrans('', '', string.punctuation)).lower() best_match = None highest_score = 0 # 遍历数据集,计算每个问题与用户输入的匹配度 for item in qa_dataset: match_score = fuzz.ratio(item['ProcessedQuestion'], processed_input) # 保留分数最高且超过阈值的匹配项 if match_score > highest_score and match_score >= threshold: highest_score = match_score best_match = item return best_match if best_match else {"error": "没有找到匹配的问题哦"}
测试模糊匹配
user_query = "伦敦在哪个国家" print(find_fuzzy_matching_answer(user_query)) # 输出:{'QuestionID': 'Q1', 'OriginalQuestion': '伦敦在哪里?', 'ProcessedQuestion': '伦敦在哪里', 'Answer': '在英国', 'Document': 'Google'}
这样处理之后,不管用户输入的格式有小差异,还是表述略有不同,都能找到对应的问答内容啦!
内容的提问来源于stack exchange,提问作者acomputerman123
相关产品推荐
相关产品推荐

