如何在R中按组移除多列含2个及以上异常值的样本
问题描述
从数据框x中提取出的数据框df包含A、B、C三组学生的四次考试成绩,具体数据如下:
| Group | ExamScore1 | ExamScore2 | ExamScore3 | ExamScore4 |
|---|---|---|---|---|
| A | 68 | 84 | 19 | 95 |
| B | 68 | 83 | 28 | 92 |
| B | 68 | 92 | 38 | 83 |
| C | 78 | 84 | 38 | 94 |
| C | 94 | 85 | 28 | 82 |
| C | 94 | 92 | 38 | 38 |
| B | 48 | 83 | 83 | 38 |
| B | 38 | 19 | 48 | 29 |
| C | 29 | 23 | 91 | 12 |
| A | 48 | 34 | 92 | 39 |
| A | 95 | 58 | 93 | 48 |
需要完成以下操作:
- 按A、B、C组分别使用四分位距法识别各考试成绩中的异常值
- 移除那些在2门及以上考试中出现异常值的学生
- 将处理后的数据存入新数据框
df2
用户提供了参考代码片段:
df1 <- df %>% group_by(x.Group) %>% filter(!x.score %in% boxplot.stats(x.score)$out) %>% ungroup()
解决方案
使用tidyverse工具包可以高效完成需求,以下是分步实现和整合代码:
分步实现
1. 转换数据格式并生成学生标识
原数据是宽格式,先转为长格式,同时生成临时student_id(原数据无唯一学生标识,用行号替代):
library(tidyverse) df_long <- df %>% mutate(student_id = row_number()) %>% pivot_longer( cols = starts_with("ExamScore"), names_to = "exam", values_to = "score" )
2. 按组标记异常值
按Group和考试科目分组,用boxplot.stats()识别异常值并标记:
df_outlier_marked <- df_long %>% group_by(Group, exam) %>% mutate(is_outlier = score %in% boxplot.stats(score)$out) %>% ungroup()
3. 过滤异常次数超标的学生
统计每个学生的异常考试次数,保留异常次数少于2次的学生,再转回宽格式:
# 统计每个学生的异常次数 student_outlier_count <- df_outlier_marked %>% group_by(student_id) %>% summarise(outlier_count = sum(is_outlier)) %>% ungroup() # 筛选并转换回宽格式 df2 <- df_outlier_marked %>% inner_join(student_outlier_count %>% filter(outlier_count < 2), by = "student_id") %>% pivot_wider(names_from = exam, values_from = score) %>% select(-student_id)
整合版代码
如果需要更简洁的一步式代码:
library(tidyverse) df2 <- df %>% mutate(student_id = row_number()) %>% pivot_longer(cols = starts_with("ExamScore"), names_to = "exam", values_to = "score") %>% group_by(Group, exam) %>% mutate(is_outlier = score %in% boxplot.stats(score)$out) %>% ungroup() %>% group_by(student_id) %>% filter(sum(is_outlier) < 2) %>% ungroup() %>% pivot_wider(names_from = exam, values_from = score) %>% select(-student_id)
关键说明
- 临时
student_id用于跟踪每个学生的跨考试异常情况,处理完成后移除 boxplot.stats()默认采用1.5倍四分位距作为异常值判定阈值,符合常规四分位距法逻辑- 最终
df2保留的是所有考试中异常次数不足2次的学生的原始成绩,并未删除单个异常值,而是直接移除符合条件的学生
内容的提问来源于stack exchange,提问作者Dan Nguyen
相关产品推荐
相关产品推荐

