如何使用scikit-learn构建基于机器学习的问答聊天机器人
Hey there! Let's break down how you can build a better Q&A chatbot using scikit-learn with your existing dataset—ditching the limitations of basic NLTK keyword matching. Here's a practical, step-by-step approach tailored to your needs:
First, make sure your data is organized into two paired lists (or a pandas DataFrame):
- A list of questions (the user inputs your bot will match against)
- A corresponding list of answers (the responses your bot should return)
For example:
questions = [ "How do I reset my password?", "Where can I find my order history?", "What's your return policy?" ] answers = [ "You can reset your password by clicking 'Forgot Password' on the login page.", "Order history is located under the 'My Account' tab in the main menu.", "We accept returns within 30 days of purchase for unused items." ]
Basic keyword matching misses nuance—let's add preprocessing to capture semantic meaning, using NLTK alongside scikit-learn:
import nltk from nltk.corpus import stopwords from nltk.stem import WordNetLemmatizer from sklearn.feature_extraction.text import TfidfVectorizer # Download required NLTK resources (run once) nltk.download('stopwords') nltk.download('wordnet') nltk.download('punkt') # Preprocessing function to clean text def preprocess_text(text): # Lowercase all text text = text.lower() # Split into individual words tokens = nltk.word_tokenize(text) # Remove stopwords (like "the", "and") and lemmatize words (e.g., "running" → "run") stop_words = set(stopwords.words('english')) lemmatizer = WordNetLemmatizer() cleaned_tokens = [ lemmatizer.lemmatize(token) for token in tokens if token not in stop_words and token.isalpha() ] return ' '.join(cleaned_tokens) # Apply preprocessing to your question bank processed_questions = [preprocess_text(q) for q in questions]
The core of your Q&A bot will be finding the most semantically similar question to a user's query. We'll use TF-IDF to convert text into numerical features, then a Nearest Neighbors model to find matches:
from sklearn.neighbors import NearestNeighbors # Convert preprocessed questions to TF-IDF vectors (captures word importance) vectorizer = TfidfVectorizer() question_vectors = vectorizer.fit_transform(processed_questions) # Train a KNN model to find the closest matching question knn_model = NearestNeighbors(n_neighbors=1, metric='cosine') knn_model.fit(question_vectors) # Function to generate answers for new queries def get_bot_response(user_query): # Preprocess the user's query processed_query = preprocess_text(user_query) # Convert query to TF-IDF vector query_vector = vectorizer.transform([processed_query]) # Find the closest matching question in your dataset distances, indices = knn_model.kneighbors(query_vector) # Return the corresponding answer return answers[indices[0][0]]
Try it out with queries that aren't exact matches to your training questions—you'll notice it handles paraphrases better than keyword matching:
# Example test queries print(get_bot_response("How do I change my password?")) # Matches "How do I reset my password?" print(get_bot_response("Can I send back an item I bought?")) # Matches "What's your return policy?"
To improve further:
- Add more training data (the more diverse your questions, the better the model)
- Experiment with different metrics (e.g.,
euclideaninstead ofcosine) - If you have categorized questions, try a classifier like
LogisticRegressioninstead of KNN to map queries to answer categories
内容的提问来源于stack exchange,提问作者Yann Droy

