You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中批量计算各Key与距离的Spearman秩相关并绘图?

Solution: Batch Spearman Rank Correlation Calculation & Visualization with Pandas

Got it, let's break this down into actionable steps. You've got a DataFrame where each column (key) has values paired with a fixed set of distances, and you want to compute Spearman rank correlation for each column against those distances, then plot the results. Here's how to do it smoothly:

Step 1: Set Up Your Data

First, let's formalize your sample data into a pandas DataFrame, and define your distance array (note: your sample arrays have 5 values, so I'll use a 5-element distance list to match—adjust if your actual data includes 0 as the sixth distance):

import pandas as pd
from scipy.stats import spearmanr
import matplotlib.pyplot as plt

# Define your fixed distance list (match length to your key arrays)
distances = [1000, 800, 600, 400, 200]

# Create sample DataFrame from your example data
data = {
    'key1': [1.21, 0.99, 6.66, 5.22, 3.33],
    'key2': [2.21, 2.99, 5.66, 6.22, 2.33],
    'key3': [4.21, 1.59, 6.66, 9.12, 0.23]
}
df = pd.DataFrame(data)

Step 2: Batch Calculate Spearman Rank Correlation

We'll use df.apply() to run the Spearman correlation calculation on each column. The spearmanr function returns a tuple of (correlation coefficient, p-value)—we'll extract just the coefficient for our core results:

# Helper function to compute Spearman correlation against the fixed distance list
def calculate_spearman(col):
    corr, p_val = spearmanr(col, distances)
    return corr

# Apply the function to every column in the DataFrame
spearman_results = df.apply(calculate_spearman).reset_index()
spearman_results.columns = ['Key', 'Spearman_Correlation']

# Print the results for quick inspection
print(spearman_results)

This gives you a clean, structured DataFrame with each key and its corresponding Spearman rank correlation to the distance sequence.

Step 3: Visualize the Results

Let's plot the correlation values to easily compare how each key relates to the distances. A bar plot works great for clear side-by-side comparison:

plt.figure(figsize=(10, 6))
plt.bar(spearman_results['Key'], spearman_results['Spearman_Correlation'], color='skyblue')
plt.axhline(y=0, color='gray', linestyle='--', alpha=0.7)
plt.title('Spearman Rank Correlation Between Keys and Distances')
plt.xlabel('Keys')
plt.ylabel('Spearman Correlation Coefficient')
plt.grid(axis='y', linestyle=':', alpha=0.6)
plt.tight_layout()
plt.show()

If your keys follow an ordered sequence, you can swap plt.bar() for plt.plot() to create a line plot that highlights trends across keys.

Quick Notes

  • If your distance list includes 0 (making it 6 elements), just update the distances array to match the length of your key arrays—pandas handles alignment automatically.
  • To include p-values in your results, modify the helper function to return both the correlation and p-value, then adjust the results DataFrame to add a p_value column.
  • This approach is efficient even for large numbers of keys, as apply() vectorizes operations across columns without manual loops.

内容的提问来源于stack exchange,提问作者Alex Trevylan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:32:56