如何用R语言将Ximputed数据中的国家归类到大洲?求助
将数据中的国家匹配到大洲的解决方案
原代码的问题
rename.vars不属于dplyr生态,语法错误;即便用gdata包的同名函数,也容易因参数格式问题出错,建议用dplyr::rename替代。- 直接用
Ximputed$country作为ne_countries的country参数,大概率因国家名称格式不匹配(比如简称、拼写差异)导致无法正确筛选国家,最终连接失败。 - 原代码未直接提取
ne_countries自带的continent字段,反而绕了弯路。
可靠解决方案(推荐用countrycode处理名称匹配)
countrycode包能自动处理绝大多数国家名称的格式差异,是解决这类匹配问题的最优工具:
- 安装并加载所需包:
install.packages(c("rnaturalearth", "sf", "dplyr", "countrycode")) library(rnaturalearth) library(sf) library(dplyr) library(countrycode)
- 获取带大洲信息的国家基础数据:
# 提取主权国家名称和大洲字段,并重命名便于后续连接 world_continent <- ne_countries(returnclass = "sf") %>% select(country = sovereignt, continent)
- 清洗数据中的国家名称,统一匹配格式:
Ximputed <- Ximputed %>% mutate( # 将原始国家名称转换为标准名称,匹配world_continent的country字段 country_standard = countrycode(country, origin = "country.name", destination = "country.name") ) # 可选:检查转换失败的条目,手动修正特殊名称 # filter(Ximputed, is.na(country_standard))
- 执行左连接,匹配大洲信息:
x2 <- Ximputed %>% left_join(world_continent, by = c("country_standard" = "country")) %>% # 可选:移除中间清洗字段 select(-country_standard)
替代方案(手动处理名称匹配,适合小数据集)
如果不想用countrycode,可以先排查名称差异再手动修正:
# 找出Ximputed中无法匹配的国家名称 unmatched_countries <- setdiff(Ximputed$country, world_continent$country) print(unmatched_countries) # 手动修正这些名称,例如: Ximputed <- Ximputed %>% mutate(country = case_when( country == "USA" ~ "United States of America", country == "UK" ~ "United Kingdom", TRUE ~ country )) # 再执行连接 x2 <- Ximputed %>% left_join(world_continent, by = "country")
内容的提问来源于stack exchange,提问作者Malle
相关产品推荐
相关产品推荐

