求助:计算Sales与Advertise的Spearman相关系数(数据拆分困惑)
Hey there! Let's get this sorted out for you—you don't need to split variables into 5 parts at all to calculate the Spearman correlation between Sales and Advertise. That's a common misunderstanding, but Spearman's rank correlation only requires the two variables you care about: their raw values (or more precisely, their rank orders).
First, let's clean up your dataset into a readable table so we're on the same page (I noticed a potential data entry issue starting from 1994—more on that later):
| Row | Year | Sales | Advertise | Employees |
|---|---|---|---|---|
| 1 | 1985 | 1.05 | 162 | 32 |
| 2 | 1986 | 1.26 | 285 | 47 |
| 3 | 1987 | 1.47 | 540 | 23 |
| 4 | 1988 | 2.16 | 261 | 68 |
| 5 | 1989 | 1.95 | 360 | 32 |
| 6 | 1990 | 2.4 | 690 | 17 |
| 7 | 1991 | 2.37 | 495 | 58 |
| 8 | 1992 | 3.15 | 948 | 75 |
| 9 | 1993 | 3.57 | 720 | 98 |
| 10 | 1994 | 4.41 | 1.14 | 43 |
| 11 | 1995 | 4.5 | 1.395 | 76 |
| 12 | 1996 | 5.61 | 1.56 | 89 |
| 13 | 1997 | 5.19 | 1.38 | 108 |
| 14 | 1998 | 5.67 | 1.26 | 76 |
| 15 | 1999 | 5.16 | 1.71 | 65 |
| 16 | 2000 | 6.84 | 1.86 | 93 |
First, a quick reality check:
Spearman correlation measures the strength and direction of the monotonic relationship between two variables. It works by looking at the rank (order) of each value in the dataset, not the actual values themselves. You don't need to split any variables here—just isolate the Sales and Advertise columns.
Critical note about your data:
Looking at the Advertise values starting from 1994, they drop from 720 to 1.14, then stay around 1-2. That seems like a possible typo (maybe you missed a zero? Like 1140 instead of 1.14?). This will drastically skew your correlation results, so double-check those numbers first!
How to calculate it with code:
I'll show you two common tools—Python and R—to compute this in seconds.
Python (using scipy):
We'll use the spearmanr function from scipy.stats. Here's the full code:
from scipy.stats import spearmanr # List out your Sales data sales = [1.05, 1.26, 1.47, 2.16, 1.95, 2.4, 2.37, 3.15, 3.57, 4.41, 4.5, 5.61, 5.19, 5.67, 5.16, 6.84] # Advertise data as provided—remember to fix the possible typos! advertise = [162, 285, 540, 261, 360, 690, 495, 948, 720, 1.14, 1.395, 1.56, 1.38, 1.26, 1.71, 1.86] # Calculate Spearman correlation and p-value corr_coeff, p_val = spearmanr(sales, advertise) print(f"Spearman Correlation Coefficient: {corr_coeff:.4f}") print(f"P-value: {p_val:.4f}")
R (using base R):
If you prefer R, use the cor() function with method = "spearman":
# Define your vectors sales <- c(1.05, 1.26, 1.47, 2.16, 1.95, 2.4, 2.37, 3.15, 3.57, 4.41, 4.5, 5.61, 5.19, 5.67, 5.16, 6.84) advertise <- c(162, 285, 540, 261, 360, 690, 495, 948, 720, 1.14, 1.395, 1.56, 1.38, 1.26, 1.71, 1.86) # Compute Spearman correlation spearman_result <- cor(sales, advertise, method = "spearman") cat("Spearman Correlation Coefficient:", round(spearman_result, 4), "\n")
What to do next:
- Fix the Advertise values from 1994 onward if they're typos.
- Run the code above—you'll get your correlation coefficient (ranging from -1 to 1) and a p-value (to test if the correlation is statistically significant).
That's it—no variable splitting needed, just clean data and a few lines of code!
内容的提问来源于stack exchange,提问作者Skilled Potato

