R语言中不同维度数据框的排名相关性检验方法
Got it, let's work through this problem step by step! The core issue here is that your two data frames have different sets of counties—so first we need to align them to only include counties present in both datasets before we can run a valid rank correlation test. Here's how to do it in R:
Step 1: Merge Data Frames to Keep Only Common Counties
We'll use an inner join to retain only counties that exist in both Cali_Income and Cali_Asthma_Rates. This ensures we have paired rank data for every observation we test.
Option 1: Base R
# Merge using base R's merge() function (all = FALSE means inner join) merged_counties <- merge(Cali_Income, Cali_Asthma_Rates, by = "County", all = FALSE) # Rename the rank columns for clarity (optional but recommended) colnames(merged_counties)[colnames(merged_counties) == "Rank.x"] <- "Income_Rank" colnames(merged_counties)[colnames(merged_counties) == "Rank.y"] <- "Asthma_Rank"
Option 2: Tidyverse (dplyr)
If you prefer the tidyverse syntax:
library(dplyr) # Inner join + rename ranks in one step merged_counties <- inner_join( Cali_Income %>% rename(Income_Rank = Rank), Cali_Asthma_Rates %>% rename(Asthma_Rank = Rank), by = "County" )
Step 2: Verify the Merged Data
Quickly check that the merged data looks correct (only shared counties, no missing ranks):
# View first few rows head(merged_counties) # Check for missing values in rank columns sum(is.na(merged_counties$Income_Rank) | is.na(merged_counties$Asthma_Rank))
Step 3: Run Spearman Rank Correlation Test
Since we're working with rank data (not continuous numerical values), the Spearman correlation test is the right choice—it measures the monotonic relationship between two rank variables, which is exactly what you need here.
# Perform the Spearman rank correlation test rank_correlation <- cor.test(merged_counties$Income_Rank, merged_counties$Asthma_Rank, method = "spearman") # Print the full result print(rank_correlation)
How to Interpret the Result:
- The
rhovalue: Ranges from -1 to 1. A positive rho means higher income rank (richer counties) correlates with higher asthma rank (more diagnoses), while a negative rho means richer counties have lower asthma ranks. - The
p-value: If p < 0.05, the correlation is statistically significant.
Key Notes:
- We can't include counties that only exist in one dataset—without paired rank data, there's no way to include them in a correlation test. This is standard practice for paired analyses.
- Spearman is preferred over Pearson here because Pearson assumes normally distributed continuous data, which ranks are not.
内容的提问来源于stack exchange,提问作者Melissa

