如何在DataFrame中创建包含美国州所属地区的新列?
为美国州缩写匹配地区的解决方案
一、利用R内置数据(无需额外安装包)
R自带了州缩写与地区的映射数据,直接结合dplyr::mutate就能快速实现:
state.abb是内置的州缩写向量,state.region是对应的地区分类(包含Northeast、South、North Central、West)- 代码示例:
library(dplyr) # 假设你的数据框是df,州缩写列名为state_abbr df <- df %>% mutate(region = state.region[match(state_abbr, state.abb)]) # 若需要将"North Central"统一为"Midwest",可追加recode操作 df <- df %>% mutate(region = recode(state.region[match(state_abbr, state.abb)], "North Central" = "Midwest"))
二、手动创建映射表(自定义划分)
如果内置分类不符合需求,或者需要自定义地区规则,手动构建映射表后关联更灵活:
- 构建州-地区映射表:
state_region_map <- tibble( state_abbr = c("IA", "IL", "IN", "NY", "MA", "TX", "CA"), region = c("Midwest", "Midwest", "Midwest", "Northeast", "Northeast", "South", "West") # 补充所有需要覆盖的州缩写及对应地区 )
- 通过
left_join关联到原数据:
df <- df %>% left_join(state_region_map, by = "state_abbr")
三、使用专用地理工具包(usmap)
如果需要更全面的地理数据支持,usmap包提供了标准的州-地区映射:
# 安装并加载包 install.packages("usmap") library(usmap) # 获取州-地区映射数据并关联 state_info <- usmap::statepop %>% select(abbr, region) %>% distinct() df <- df %>% left_join(state_info, by = c("state_abbr" = "abbr"))
内容的提问来源于stack exchange,提问作者piper180
相关产品推荐
相关产品推荐

