You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中不同维度数据框的排名相关性检验方法

Solution for Testing Rank Correlation Between Mismatched County Datasets

Got it, let's work through this problem step by step! The core issue here is that your two data frames have different sets of counties—so first we need to align them to only include counties present in both datasets before we can run a valid rank correlation test. Here's how to do it in R:

Step 1: Merge Data Frames to Keep Only Common Counties

We'll use an inner join to retain only counties that exist in both Cali_Income and Cali_Asthma_Rates. This ensures we have paired rank data for every observation we test.

Option 1: Base R

# Merge using base R's merge() function (all = FALSE means inner join)
merged_counties <- merge(Cali_Income, Cali_Asthma_Rates, by = "County", all = FALSE)

# Rename the rank columns for clarity (optional but recommended)
colnames(merged_counties)[colnames(merged_counties) == "Rank.x"] <- "Income_Rank"
colnames(merged_counties)[colnames(merged_counties) == "Rank.y"] <- "Asthma_Rank"

Option 2: Tidyverse (dplyr)

If you prefer the tidyverse syntax:

library(dplyr)

# Inner join + rename ranks in one step
merged_counties <- inner_join(
  Cali_Income %>% rename(Income_Rank = Rank),
  Cali_Asthma_Rates %>% rename(Asthma_Rank = Rank),
  by = "County"
)

Step 2: Verify the Merged Data

Quickly check that the merged data looks correct (only shared counties, no missing ranks):

# View first few rows
head(merged_counties)

# Check for missing values in rank columns
sum(is.na(merged_counties$Income_Rank) | is.na(merged_counties$Asthma_Rank))

Step 3: Run Spearman Rank Correlation Test

Since we're working with rank data (not continuous numerical values), the Spearman correlation test is the right choice—it measures the monotonic relationship between two rank variables, which is exactly what you need here.

# Perform the Spearman rank correlation test
rank_correlation <- cor.test(merged_counties$Income_Rank, merged_counties$Asthma_Rank, method = "spearman")

# Print the full result
print(rank_correlation)

How to Interpret the Result:

  • The rho value: Ranges from -1 to 1. A positive rho means higher income rank (richer counties) correlates with higher asthma rank (more diagnoses), while a negative rho means richer counties have lower asthma ranks.
  • The p-value: If p < 0.05, the correlation is statistically significant.

Key Notes:

  • We can't include counties that only exist in one dataset—without paired rank data, there's no way to include them in a correlation test. This is standard practice for paired analyses.
  • Spearman is preferred over Pearson here because Pearson assumes normally distributed continuous data, which ranks are not.

内容的提问来源于stack exchange,提问作者Melissa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:07:46