You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:计算Sales与Advertise的Spearman相关系数(数据拆分困惑)

Hey there! Let's get this sorted out for you—you don't need to split variables into 5 parts at all to calculate the Spearman correlation between Sales and Advertise. That's a common misunderstanding, but Spearman's rank correlation only requires the two variables you care about: their raw values (or more precisely, their rank orders).

First, let's clean up your dataset into a readable table so we're on the same page (I noticed a potential data entry issue starting from 1994—more on that later):

RowYearSalesAdvertiseEmployees
119851.0516232
219861.2628547
319871.4754023
419882.1626168
519891.9536032
619902.469017
719912.3749558
819923.1594875
919933.5772098
1019944.411.1443
1119954.51.39576
1219965.611.5689
1319975.191.38108
1419985.671.2676
1519995.161.7165
1620006.841.8693

First, a quick reality check:

Spearman correlation measures the strength and direction of the monotonic relationship between two variables. It works by looking at the rank (order) of each value in the dataset, not the actual values themselves. You don't need to split any variables here—just isolate the Sales and Advertise columns.

Critical note about your data:

Looking at the Advertise values starting from 1994, they drop from 720 to 1.14, then stay around 1-2. That seems like a possible typo (maybe you missed a zero? Like 1140 instead of 1.14?). This will drastically skew your correlation results, so double-check those numbers first!

How to calculate it with code:

I'll show you two common tools—Python and R—to compute this in seconds.

Python (using scipy):

We'll use the spearmanr function from scipy.stats. Here's the full code:

from scipy.stats import spearmanr

# List out your Sales data
sales = [1.05, 1.26, 1.47, 2.16, 1.95, 2.4, 2.37, 3.15, 3.57, 4.41, 4.5, 5.61, 5.19, 5.67, 5.16, 6.84]
# Advertise data as provided—remember to fix the possible typos!
advertise = [162, 285, 540, 261, 360, 690, 495, 948, 720, 1.14, 1.395, 1.56, 1.38, 1.26, 1.71, 1.86]

# Calculate Spearman correlation and p-value
corr_coeff, p_val = spearmanr(sales, advertise)

print(f"Spearman Correlation Coefficient: {corr_coeff:.4f}")
print(f"P-value: {p_val:.4f}")

R (using base R):

If you prefer R, use the cor() function with method = "spearman":

# Define your vectors
sales <- c(1.05, 1.26, 1.47, 2.16, 1.95, 2.4, 2.37, 3.15, 3.57, 4.41, 4.5, 5.61, 5.19, 5.67, 5.16, 6.84)
advertise <- c(162, 285, 540, 261, 360, 690, 495, 948, 720, 1.14, 1.395, 1.56, 1.38, 1.26, 1.71, 1.86)

# Compute Spearman correlation
spearman_result <- cor(sales, advertise, method = "spearman")

cat("Spearman Correlation Coefficient:", round(spearman_result, 4), "\n")

What to do next:

  1. Fix the Advertise values from 1994 onward if they're typos.
  2. Run the code above—you'll get your correlation coefficient (ranging from -1 to 1) and a p-value (to test if the correlation is statistically significant).

That's it—no variable splitting needed, just clean data and a few lines of code!

内容的提问来源于stack exchange,提问作者Skilled Potato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:04:43