如何在R中用country_codes数据框替换test_data的国家代码列值?
解决方案
这里提供两种常用的实现方式,都能完成将国家代码替换为对应ID的需求:
方法1:基础R原生实现
利用match()函数匹配元素位置,直接替换目标列:
# 替换origin_country列 test_data$origin_country <- as.character(country_codes$ID[match(test_data$origin_country, country_codes$countries)]) # 替换destin_country列 test_data$destin_country <- as.character(country_codes$ID[match(test_data$destin_country, country_codes$countries)]) # 查看最终结果 print(test_data)
说明:match(x, table)返回x中每个元素在table中的索引位置,通过该索引提取country_codes$ID对应的数值,最后用as.character()转换为字符型,与你期望的输出格式一致。
方法2:dplyr包实现(适合复杂数据场景)
借助dplyr的连接操作,更直观地完成匹配替换:
library(dplyr) output <- test_data %>% # 匹配origin_country对应的ID left_join(country_codes, by = c("origin_country" = "countries")) %>% rename(origin_id = ID) %>% # 匹配destin_country对应的ID left_join(country_codes, by = c("destin_country" = "countries")) %>% rename(destin_id = ID) %>% # 调整列顺序与名称 select(origin_country = origin_id, destin_country = destin_id, year, item, value) %>% # 转换为字符型 mutate(across(c(origin_country, destin_country), as.character)) # 查看最终结果 print(output)
说明:通过两次left_join分别关联出发地和目的地的ID,再通过rename()和select()调整列名与顺序,最后统一转换为字符型。
内容的提问来源于stack exchange,提问作者Grecchi
相关产品推荐
相关产品推荐

