You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python NLTK实现用户输入与Excel问题数据的相似度匹配?

解决Python聊天机器人的问题匹配需求

嘿,我来帮你搞定这个问题!你已经成功把CSV数据读进来了,接下来核心就是问题匹配——既要处理用户输入的小差异(比如有没有问号、大小写不同),又要准确找到对应的问答条目。下面分步骤给你讲清楚怎么做:

第一步:优化数据存储与预处理

首先,我们需要把读取到的数据做标准化处理,这样能消除用户输入的格式差异。比如把问题里的标点去掉、转成小写,这样不管用户输入“伦敦在哪里?”还是“伦敦在哪里”,处理后都是一样的内容,方便匹配。

修改后的读取代码如下:

import csv
import string

# 读取CSV数据,存储为包含完整信息的列表(包含预处理后的问题)
qa_dataset = []
with open("dataset.csv", encoding="utf-8") as csvfile:
    reader = csv.DictReader(csvfile)
    for row in reader:
        # 预处理问题:移除所有标点、转成小写
        processed_question = row['Question'].translate(str.maketrans('', '', string.punctuation)).lower()
        qa_dataset.append({
            'QuestionID': row['QuestionID'],
            'OriginalQuestion': row['Question'],
            'ProcessedQuestion': processed_question,
            'Answer': row['Answer'],
            'Document': row['Document']
        })

第二步:实现精确匹配函数

接下来写一个匹配函数,接收用户输入后,用同样的规则预处理,然后遍历数据集找到完全匹配的条目,返回完整数据:

def find_matching_answer(user_input):
    # 预处理用户输入,和存储的问题保持一致规则
    processed_input = user_input.translate(str.maketrans('', '', string.punctuation)).lower()
    # 遍历数据集查找匹配项
    for item in qa_dataset:
        if item['ProcessedQuestion'] == processed_input:
            return item  # 返回包含QuestionID、Answer、Document的完整条目
    # 没有找到匹配时返回提示
    return {"error": "没有找到对应的问题哦"}

测试一下效果

# 测试带问号的输入
user_query1 = "伦敦在哪里?"
print(find_matching_answer(user_query1))
# 输出:{'QuestionID': 'Q1', 'OriginalQuestion': '伦敦在哪里?', 'ProcessedQuestion': '伦敦在哪里', 'Answer': '在英国', 'Document': 'Google'}

# 测试不带问号的输入
user_query2 = "伦敦在哪里"
print(find_matching_answer(user_query2))
# 输出和上面完全一致

# 测试另一个问题
user_query3 = "球场上有多少足球运动员?"
print(find_matching_answer(user_query3))
# 输出:{'QuestionID': 'Q2', 'OriginalQuestion': '球场上有多少足球运动员?', 'ProcessedQuestion': '球场上有多少足球运动员', 'Answer': '22', 'Document': 'Google'}

第三步:扩展模糊匹配(可选)

如果需要处理用户表述不同但意思相近的问题(比如用户问“伦敦在哪个国家”,想匹配到“伦敦在哪里?”),可以用模糊匹配库fuzzywuzzy来实现:

  1. 先安装依赖:
pip install fuzzywuzzy python-Levenshtein
  1. 编写模糊匹配函数:
from fuzzywuzzy import fuzz

def find_fuzzy_matching_answer(user_input, threshold=80):
    processed_input = user_input.translate(str.maketrans('', '', string.punctuation)).lower()
    best_match = None
    highest_score = 0
    # 遍历数据集,计算每个问题与用户输入的匹配度
    for item in qa_dataset:
        match_score = fuzz.ratio(item['ProcessedQuestion'], processed_input)
        # 保留分数最高且超过阈值的匹配项
        if match_score > highest_score and match_score >= threshold:
            highest_score = match_score
            best_match = item
    return best_match if best_match else {"error": "没有找到匹配的问题哦"}

测试模糊匹配

user_query = "伦敦在哪个国家"
print(find_fuzzy_matching_answer(user_query))
# 输出:{'QuestionID': 'Q1', 'OriginalQuestion': '伦敦在哪里?', 'ProcessedQuestion': '伦敦在哪里', 'Answer': '在英国', 'Document': 'Google'}

这样处理之后,不管用户输入的格式有小差异,还是表述略有不同,都能找到对应的问答内容啦!

内容的提问来源于stack exchange,提问作者acomputerman123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:43:14