如何计算两个Pandas DataFrame对应行的余弦相似度(非全成对cdist)
The issue with using cdist here is that it calculates the pairwise cosine distance between all rows of the two DataFrames, resulting in a full matrix. To get just the similarity between corresponding rows (i.e., row 0 of trigger with row 0 of action, row 1 with row 1, etc.), you have two main approaches—one that's memory-efficient and fast, and another that repurposes your existing cdist call.
Method 1: Vectorized Numpy Calculation (Recommended)
This approach computes the cosine similarity directly for each pair of corresponding rows without generating the full pairwise matrix, which saves significant memory and computation time.
Cosine similarity between two vectors (u) and (v) is defined as:
[
\text{cos_sim}(u, v) = \frac{u \cdot v}{||u|| \times ||v||}
]
Here's how to implement this with vectorized operations:
import numpy as np # Convert DataFrames to numpy arrays (if not already) trigger_arr = trigger.values action_arr = action.values # Calculate dot product for each corresponding row pair dot_products = (trigger_arr * action_arr).sum(axis=1) # Calculate L2 norm for each row in both DataFrames trigger_norms = np.linalg.norm(trigger_arr, axis=1) action_norms = np.linalg.norm(action_arr, axis=1) # Compute cosine similarity for each row pair cosine_similarities = dot_products / (trigger_norms * action_norms) # Reshape to (20675, 1) as desired result = cosine_similarities.reshape(-1, 1)
Method 2: Extract Diagonal from cdist Result
If you want to stick with scipy.spatial.distance.cdist, you can extract the diagonal elements of the resulting matrix (since those correspond to row-wise pairs). Note that cdist returns cosine distance (which is (1 - \text{cosine similarity})), so we need to adjust for that:
from scipy.spatial.distance import cdist import numpy as np # Compute full cosine distance matrix cosine_dist_matrix = cdist(trigger, action, metric='cosine') # Extract diagonal (row-wise pairs) and convert to similarity cosine_similarities = 1 - np.diag(cosine_dist_matrix) # Reshape to (20675, 1) result = cosine_similarities.reshape(-1, 1)
Important Note on Method 2
This method generates a (20675 \times 20675) matrix, which uses around 3.2GB of memory (for 64-bit floats). For large datasets like yours, Method 1 is far more efficient and avoids unnecessary memory usage.
内容的提问来源于stack exchange,提问作者profhoff

