You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算DataFrame行级余弦相似度?含首行与其余行对比需求

Hey there! Let's break down how to tackle both of your DataFrame cosine similarity questions step by step. We'll use the sklearn.metrics.pairwise.cosine_similarity function since it's straightforward and widely used for this task—no need to reinvent the wheel!

1. Calculating Row-Level Cosine Similarity for a Python DataFrame

First off, cosine similarity measures the cosine of the angle between two vectors, which tells us how similar they are in direction. For a DataFrame, each row acts as a vector we want to compare against every other row.

Key Prep Notes:

  • Make sure your DataFrame only contains numerical columns—cosine similarity doesn't work with text or categorical data. Drop or encode non-numeric columns first if needed.
  • The cosine_similarity function returns a square matrix where the value at position (i,j) is the similarity score between row i and row j.

Example Code:

import pandas as pd
from sklearn.metrics.pairwise import cosine_similarity

# Create a sample numerical DataFrame
df = pd.DataFrame({
    'col1': [1, 2, 3, 4],
    'col2': [5, 6, 7, 8],
    'col3': [9, 10, 11, 12]
})

# Calculate row-level cosine similarity
similarity_matrix = cosine_similarity(df)

# Convert to a labeled DataFrame for readability
similarity_df = pd.DataFrame(similarity_matrix, index=df.index, columns=df.index)
print(similarity_df)

Output Breakdown:

The resulting similarity_df will show that each row has a perfect similarity score of 1.0 with itself. Values closer to 1 mean the rows are more similar in direction—for our sample data, all rows are highly similar since they follow a linear pattern.

2. Calculating Cosine Similarity Between the First Row and All Remaining Rows

This is just a targeted subset of the full similarity matrix. You can either pull the relevant values from the full matrix, or compute it directly to save computation time (useful for large DataFrames).

Method 1: Reuse the Full Similarity Matrix

# Extract similarity scores between the first row (index 0) and all other rows
first_row_similarities = similarity_matrix[0]

# Convert to a labeled Series for clarity
similarity_series = pd.Series(first_row_similarities, index=df.index, name='Similarity to Row 0')
print(similarity_series)

Method 2: Direct Computation (More Efficient)

Skip calculating the full matrix and only compare the first row to the rest:

# Extract the first row as a 2D array (required by cosine_similarity)
first_row = df.iloc[[0]]

# Compute similarity between the first row and all rows in the DataFrame
first_row_similarities = cosine_similarity(first_row, df)[0]

# Convert to a labeled Series
similarity_series = pd.Series(first_row_similarities, index=df.index, name='Similarity to Row 0')
print(similarity_series)

Output Breakdown:

You'll get a Series where the first value is 1.0 (perfect similarity to itself), followed by scores that show how similar each remaining row is to the first one.

内容的提问来源于stack exchange,提问作者Abhi sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:46:06