基于SVD隐特征的物品相似度计算及快速非负矩阵分解方法咨询
Awesome question—let's break this down into two clear parts since you're working with scipy.sparse.linalg.svds for your user-item rating matrix.
First, let's clarify the output of svds for your setup (rows = users, cols = items):
U:(num_users, k)matrix (each row is a user's latent feature vector)s:(k,)array of singular values (weights for each latent dimension)Vh:(k, num_items)matrix (transpose of the item latent matrixV)
To get meaningful item features, you should scale Vh by the singular values (they represent how important each latent dimension is). Here's how to get a clean (num_items, k) matrix where each row is an item's latent vector:
import numpy as np # Scale Vh by singular values and transpose to item-major order item_features = (np.diag(s) @ Vh).T # Shape: (num_items, k)
Even though these vectors have positive and negative values, you can compute valid item similarity using metrics that handle signed values:
- Cosine Similarity: The go-to choice for latent space similarity. It measures the angle between two vectors, so negative values just indicate inverse relationships (e.g., users who love Item A often hate Item B).
from scipy.spatial.distance import cosine # Get vectors for two items (replace with your item indices) vec_item1 = item_features[item_idx1] vec_item2 = item_features[item_idx2] # Cosine similarity = 1 - cosine distance similarity = 1 - cosine(vec_item1, vec_item2) # Range: [-1, 1] - Pearson Correlation: Normalizes vectors by their mean, which can reduce bias from overall latent feature magnitude.
- Dot Product: A simple unnormalized measure — higher absolute values mean stronger alignment (positive or negative) between items.
If you want to ignore negative similarity (e.g., only focus on items that appeal to similar users), you can clamp the result to 0: similarity = max(similarity, 0). Just note this discards useful signal about inverse item relationships.
If you need latent features with only non-negative values (often more interpretable, as they map to "positive" factors like "action movie" or "organic food"), here are the best options that match svds's speed:
Scikit-learn's NMF (Non-Negative Matrix Factorization): The standard non-negative alternative. It supports sparse matrices (perfect for rating data) and has a fast coordinate descent solver. For
k=50, it's roughly comparable in speed tosvdsfor most datasets.from sklearn.decomposition import NMF # Assume your user-item matrix is called `ratings_matrix` (sparse or dense) nmf = NMF(n_components=50, solver='cd', init='nndsvd', random_state=42) user_features_nmf = nmf.fit_transform(ratings_matrix) item_features_nmf = nmf.components_.T # Shape: (num_items, 50) — non-negative!The
init='nndsvd'option initializes factors using truncated SVD (with non-negative constraints), which speeds up convergence and improves result quality.Projected Gradient NMF: A variant of NMF that uses projected gradient descent for faster convergence. Scikit-learn's
NMFwithsolver='cd'(coordinate descent) is essentially this, and it's faster than the default multiplicative updates.Partial Non-Negative SVD: Some methods enforce non-negativity on one factor matrix (either user or item), but full non-negative SVD is computationally heavier than standard SVD. NMF is a better bet for speed and interpretability.
One thing to note: Unlike SVD (which has an exact solution), NMF finds a local optimum. But with good initialization (like nndsvd), you'll get stable, high-quality results quickly.
内容的提问来源于stack exchange,提问作者Pranay

