基于KNN与余弦相似度的无用户信息图书推荐实现求助
Hey there! You’re already off to a great start with building the user-item matrix and setting up your KNN model. Let’s walk through exactly how to turn that framework into personalized book recommendations for a specific user.
Step 1: Finish Training Your KNN Model
First, you need to fit the model to your user-item matrix—you’ve defined the model but haven’t fed it the data yet:
model_knn.fit(product_matrix)
Step 2: Locate the Target User in the Matrix
Pick the user you want to generate recommendations for (e.g., 45michael), then find their position in the matrix:
target_user = "45michael" user_index = df_matrix_dummy.index.get_loc(target_user) user_vector = product_matrix[user_index]
Step 3: Find the Most Similar Users
Use the kneighbors method to pull the closest users based on cosine similarity. This returns two arrays: distances to each neighbor and their indices in the matrix. We’ll skip the first result (it’s the user themselves, since they’re most similar to their own purchase history):
# Get top 6 nearest neighbors (we'll exclude the user's own entry) distances, indices = model_knn.kneighbors(user_vector, n_neighbors=6) # Isolate indices of actual similar users similar_user_indices = indices[0][1:]
Step 4: Collect Books from Similar Users
Now, gather all books that similar users have bought but the target user hasn’t. We’ll count how many times each book appears across neighbors to prioritize the most popular picks among similar buyers:
from collections import defaultdict # Get the list of books the target user already owns user_purchased = df_matrix_dummy.loc[target_user][df_matrix_dummy.loc[target_user] == 1].index # Initialize a counter to track recommendation frequency book_recommendations = defaultdict(int) # Loop through each similar user to collect their unique books for idx in similar_user_indices: similar_user = df_matrix_dummy.index[idx] # Get books the similar user purchased that the target user didn't similar_user_books = df_matrix_dummy.loc[similar_user][df_matrix_dummy.loc[similar_user] == 1].index new_books = similar_user_books.difference(user_purchased) # Count each book's occurrence for book in new_books: book_recommendations[book] += 1
Step 5: Sort and Deliver Top Recommendations
Sort the books by how many similar users bought them, then pick the top N recommendations:
# Sort books by recommendation count (descending order) sorted_recs = sorted(book_recommendations.items(), key=lambda x: x[1], reverse=True) # Get top 3 recommendations (adjust the number as needed) top_recs = [book_id for book_id, count in sorted_recs[:3]] print(f"Top book recommendations for {target_user}: {top_recs}")
Quick Edge Case Notes
- For users with no purchase history (like
7762hcin your sample matrix), you can’t use similarity-based recommendations. Instead, you could recommend the most popular books overall (count how many users bought each book and pick the top ones). - Tweak the
n_neighborsparameter based on your dataset size—more neighbors can add diversity, but too many might dilute the relevance of recommendations.
内容的提问来源于stack exchange,提问作者user6544783

