统计学期末项目:将Finance数据集edu变量分为大学/非大学两组求助
划分教育水平为大学/非大学组的解决方案
第一步:明确分组对应类别
先把edu变量按需求归类:
- 非大学组:trade school degree、some high school、less than high school、high school diploma / GED、some college, no degree
- 大学组:Graduate degree、doctorate / post graduate、bachelor's degree、associate's degree
第二步:用基础R实现(无需额外安装包)
先给数据集新增分组变量,再拆分:
# 定义非大学类别的列表 non_college_list <- c("trade school degree", "some high school", "less than high school", "high school diploma / GED", "some college, no degree") # 新增分组列edu_group Finance$edu_group <- ifelse(Finance$edu %in% non_college_list, "非大学组", "大学组") # 拆分出非大学组数据集 non_college_data <- subset(Finance, edu_group == "非大学组") # 拆分出大学组数据集 college_data <- subset(Finance, edu_group == "大学组") # 也可以用split直接生成包含两个数据集的列表 split_data <- split(Finance, Finance$edu_group) # 调用时用 split_data$`非大学组` 或 split_data$`大学组`
第三步:用dplyr实现(更直观,适合新手)
如果习惯用tidyverse工具包,代码可读性更强:
# 先安装dplyr(仅第一次用需要) # install.packages("dplyr") library(dplyr) # 生成非大学组数据集 non_college_data <- Finance %>% mutate(edu_group = case_when( edu %in% non_college_list ~ "非大学组", TRUE ~ "大学组" )) %>% filter(edu_group == "非大学组") # 生成大学组数据集 college_data <- Finance %>% mutate(edu_group = case_when( edu %in% non_college_list ~ "非大学组", TRUE ~ "大学组" )) %>% filter(edu_group == "大学组") # 一键拆分两组到列表 split_data <- Finance %>% mutate(edu_group = case_when( edu %in% non_college_list ~ "非大学组", TRUE ~ "大学组" )) %>% group_split(edu_group)
注意事项
- 先运行
unique(Finance$edu)确认数据集中edu的所有取值,确保和你列出的类别拼写、大小写、空格完全一致,否则会出现分组错误 - 如果有缺失值,
ifelse和case_when会自动把缺失值归到大学组,如果需要单独处理,可以在条件里加is.na(edu) ~ "缺失组"
内容的提问来源于stack exchange,提问作者Dalton Turner
相关产品推荐
相关产品推荐

