基于Python与LightFM的用户推荐引擎:Matrix Factorization下评分溯源问题
Great question—let’s break this down clearly since LightFM’s hybrid nature (combining matrix factorization and feature-based modeling) can make user influence a bit opaque at first. I’ll cover both your questions with practical code examples tailored to your Python/LightFM workflow.
LightFM doesn’t rely on explicit "nearest neighbor" logic like traditional user-based collaborative filtering, but its matrix factorization component learns user embeddings that encode preference similarity. The users impacting a target user’s predicted ratings fall into two core groups:
- Users with overlapping interaction behavior: If User A has interacted with (rated, purchased, clicked) similar items to your target user, their behavior helps shape the target user’s embedding during training. For example, if your target user liked Items X and Y, and User B liked X, Y, and Z, the model will learn that "users who like X/Y also like Z"—and User B’s interaction with Z directly influences the target user’s predicted rating for Z.
- Users with similar feature profiles: If you’ve included user features (e.g., age, gender, interests) in training, the model merges these into user embeddings. Users who share key features with your target user (e.g., same age group, similar hobbies) will have aligned embeddings, so their interaction patterns indirectly influence the target’s predicted ratings.
In short: any user whose interactions or features helped the model learn the target user’s preference embedding will impact their predicted scores.
To pinpoint these users, you’ll leverage the user embeddings LightFM learns during training. Here’s a step-by-step implementation:
Step 1: Extract trained user embeddings
After training your model, the user_embeddings attribute holds a matrix where each row is a user’s learned embedding vector (combining ID-based and feature-based signals if you used user features).
Step 2: Calculate similarity between target user and all others
Use cosine similarity (the standard for embedding similarity) to measure how close the target user’s embedding is to every other user’s embedding. Higher similarity means stronger influence on predictions.
Step 3: Rank and inspect top similar users
Sort users by similarity score, then validate their interactions to confirm alignment with the target user’s preferences.
Here’s the code:
from lightfm import LightFM from sklearn.metrics.pairwise import cosine_similarity import numpy as np # Assume you've already trained your model and have your interaction matrix (csr_matrix) model = LightFM(loss='warp') # Use 'logistic' for explicit ratings model.fit( interactions, user_features=user_features, # Omit if you don't use user features epochs=30, num_threads=4 ) # Define your target user ID target_user_id = 100 # Extract all user embeddings user_embeddings = model.user_embeddings # Get the target user's embedding (reshape for similarity calculation) target_embedding = user_embeddings[target_user_id].reshape(1, -1) # Compute cosine similarity between target and all other users similarity_scores = cosine_similarity(target_embedding, user_embeddings)[0] # Get top 10 most similar users (exclude the target user themselves) top_similar_indices = np.argsort(-similarity_scores)[1:11] top_similar_users = list(zip(top_similar_indices, similarity_scores[top_similar_indices])) # Print results print(f"Top 10 users influencing User {target_user_id}'s predictions:") for user_id, score in top_similar_users: print(f"User {user_id}: Similarity Score = {score:.4f}") # Optional: Verify overlapping interactions to confirm similarity def get_user_interacted_items(user_id, interaction_matrix): # Return list of item IDs the user interacted with return interaction_matrix[user_id].indices target_items = get_user_interacted_items(target_user_id, interactions) for user_id, _ in top_similar_users: similar_user_items = get_user_interacted_items(user_id, interactions) overlapping_items = np.intersect1d(target_items, similar_user_items) print(f"\nUser {user_id} shares {len(overlapping_items)} items with target user:") print(f"Sample overlapping items: {overlapping_items[:5]}..." if len(overlapping_items) > 0 else "No overlapping items")
Key Notes:
- If you used explicit ratings (instead of implicit interactions), swap the loss function to
loss='logistic'—the embedding logic and similarity calculation remain the same. - The
user_embeddingsmatrix includes both ID-based embeddings and feature-based embeddings (if provided), so similarity scores reflect both behavioral and feature-based similarity. - Users with the highest similarity scores are the ones whose interactions and features most shaped the target user’s embedding, directly impacting their predicted ratings.
内容的提问来源于stack exchange,提问作者dapo

