You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言基于borrower_country和year合并数据集填充NA失败问题

问题排查

  • 核心错误:合并时将待填充的GDP_per_capita_USD、GDP_per_capita_growth、inflation三列加入了匹配键by参数中。你已经将test数据集里日本对应的这三列全部设为NA,而gdp_infl里对应列是有效值,NA和有效值无法匹配,自然返回NA。
  • 匹配逻辑偏差:你的需求是仅按borrower_country和year两个字段匹配,用gdp_infl的数值填充test的空值,不需要把待填充列放到匹配条件里。

解决方案

方法1:修正merge参数(基础R实现)

# 仅用两个关联字段合并,不要加入待填充列
temp <- merge(test, gdp_infl, by = c("borrower_country", "year"), all.x = TRUE)

# 用gdp_infl的有效值填充test的NA
temp$GDP_per_capita_USD <- ifelse(is.na(temp$GDP_per_capita_USD.x), temp$GDP_per_capita_USD.y, temp$GDP_per_capita_USD.x)
temp$GDP_per_capita_growth <- ifelse(is.na(temp$GDP_per_capita_growth.x), temp$GDP_per_capita_growth.y, temp$GDP_per_capita_growth.x)
temp$inflation <- ifelse(is.na(temp$inflation.x), temp$inflation.y, temp$inflation.x)

# 清理多余的后缀列,保留最终需要的字段
temp <- temp[, c("borrower_country", "year", "GDP_per_capita_USD", "GDP_per_capita_growth", "inflation")]

方法2:dplyr简洁实现(适配全国家批量处理场景)

library(dplyr)
test_filled <- test %>%
  # 仅按两个关联字段左连接
  left_join(gdp_infl, by = c("borrower_country", "year"), suffix = c("", "_ref")) %>%
  # coalesce返回第一个非空值,自动用参考值填充NA
  mutate(
    GDP_per_capita_USD = coalesce(GDP_per_capita_USD, GDP_per_capita_USD_ref),
    GDP_per_capita_growth = coalesce(GDP_per_capita_growth, GDP_per_capita_growth_ref),
    inflation = coalesce(inflation, inflation_ref)
  ) %>%
  # 移除多余的参考列
  select(-ends_with("_ref"))

该逻辑可以直接复用到所有国家的填充流程,不需要单独调整日本相关代码。

内容的提问来源于stack exchange,提问作者Moz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 07:45:00