R语言基于borrower_country和year合并数据集填充NA失败问题
问题排查
- 核心错误:合并时将待填充的
GDP_per_capita_USD、GDP_per_capita_growth、inflation三列加入了匹配键by参数中。你已经将test数据集里日本对应的这三列全部设为NA,而gdp_infl里对应列是有效值,NA和有效值无法匹配,自然返回NA。 - 匹配逻辑偏差:你的需求是仅按
borrower_country和year两个字段匹配,用gdp_infl的数值填充test的空值,不需要把待填充列放到匹配条件里。
解决方案
方法1:修正merge参数(基础R实现)
# 仅用两个关联字段合并,不要加入待填充列 temp <- merge(test, gdp_infl, by = c("borrower_country", "year"), all.x = TRUE) # 用gdp_infl的有效值填充test的NA temp$GDP_per_capita_USD <- ifelse(is.na(temp$GDP_per_capita_USD.x), temp$GDP_per_capita_USD.y, temp$GDP_per_capita_USD.x) temp$GDP_per_capita_growth <- ifelse(is.na(temp$GDP_per_capita_growth.x), temp$GDP_per_capita_growth.y, temp$GDP_per_capita_growth.x) temp$inflation <- ifelse(is.na(temp$inflation.x), temp$inflation.y, temp$inflation.x) # 清理多余的后缀列,保留最终需要的字段 temp <- temp[, c("borrower_country", "year", "GDP_per_capita_USD", "GDP_per_capita_growth", "inflation")]
方法2:dplyr简洁实现(适配全国家批量处理场景)
library(dplyr) test_filled <- test %>% # 仅按两个关联字段左连接 left_join(gdp_infl, by = c("borrower_country", "year"), suffix = c("", "_ref")) %>% # coalesce返回第一个非空值,自动用参考值填充NA mutate( GDP_per_capita_USD = coalesce(GDP_per_capita_USD, GDP_per_capita_USD_ref), GDP_per_capita_growth = coalesce(GDP_per_capita_growth, GDP_per_capita_growth_ref), inflation = coalesce(inflation, inflation_ref) ) %>% # 移除多余的参考列 select(-ends_with("_ref"))
该逻辑可以直接复用到所有国家的填充流程,不需要单独调整日本相关代码。
内容的提问来源于stack exchange,提问作者Moz
相关产品推荐
相关产品推荐

