R语言含NA值列计算:两组国家成对相关性如何实现
R语言指定行范围计算两组国家两两相关性的实现
实现逻辑
首先明确匹配规则:
- 德国仅使用第3行及以后的有效数据,匹配同期美、加、英数据计算相关性
- 意大利使用全部10行数据匹配计算
- 日本仅使用第4行及以后的有效数据,匹配同期美、加、英数据计算相关性
完整实现代码(tidyverse版本)
首先加载依赖包,导入测试数据:
library(tidyverse) # 测试数据构造 df11 <- tibble( date = 2001:2010, Germany = runif(10), Italy = runif(10), Japan = runif(10), US = runif(10), Canada = runif(10), UK = runif(10) ) df11$Germany[1:2] <- NA df11$Japan[1:3] <- NA
然后定义分组和行规则,执行计算:
# 定义国家分组 front_countries <- c("Germany", "Italy", "Japan") back_countries <- c("US", "Canada", "UK") # 定义每个前组国家的计算起始行 start_row_rule <- c(Germany = 3, Italy = 1, Japan = 4) # 自定义相关性计算函数 calc_cor <- function(f_country) { start <- start_row_rule[f_country] # 截取有效行 sub_df <- df11[start:nrow(df11), ] # 计算当前前组国家与所有后组国家的相关系数 cor_res <- cor(sub_df[[f_country]], sub_df[, back_countries]) # 转为数据框返回 as_tibble(cor_res) %>% mutate(front_country = f_country, .before = 1) } # 遍历所有前组国家计算,合并结果 cor_result <- map_dfr(front_countries, calc_cor)
结果查看
直接打印cor_result即可得到结构化的相关性结果,示例输出结构如下:
| front_country | US | Canada | UK |
|---|---|---|---|
| Germany | 0.xxxx | 0.xxxx | 0.xxxx |
| Italy | 0.xxxx | 0.xxxx | 0.xxxx |
| Japan | 0.xxxx | 0.xxxx | 0.xxxx |
基础R实现版本(无需加载第三方包)
front_countries <- c("Germany", "Italy", "Japan") back_countries <- c("US", "Canada", "UK") start_row_rule <- c(Germany = 3, Italy = 1, Japan = 4) cor_result <- data.frame(front_country = front_countries) for (f in front_countries) { start <- start_row_rule[f] sub_df <- df11[start:nrow(df11), ] cor_vals <- cor(sub_df[[f]], sub_df[, back_countries]) cor_result[cor_result$front_country == f, back_countries] <- cor_vals }
内容的提问来源于stack exchange,提问作者AutumnWest
相关产品推荐
相关产品推荐

