R语言:修改数据框中重复Other的rater标签并保留至原数据
问题:修改DataFrame中重复的Other标签
我有一个包含id、rater、raterscore字段的DataFrame,每个id对应rater1、rater2、Other各一条评分,但部分id存在两个Other评分,需要把这些id的第二个Other标签改成Other2。
我试了下面的代码,但用filter后修改的内容没法保留到主数据框里:
d %>% filter(rater=="Other") %>% group_by(id) %>% filter(n()>1) %>% mutate(rater= case_when(ifelse(row_number()%%2==0,"Other2",rater))
当前数据结构
| id | rater | raterscore |
|---|---|---|
| 1 | rater2 | 3 |
| 1 | rater1 | 2 |
| 1 | Other | 3 |
| 2 | rater1 | 1.5 |
| 2 | rater2 | 3.2 |
| 2 | Other | 1.1 |
| 2 | Other | 2 |
| 3 | rater1 | 2.5 |
| 3 | rater2 | 2.7 |
| 3 | Other | 2.1 |
| 3 | Other | 2 |
| 4 | rater1 | 2.5 |
| 4 | rater2 | 2.7 |
| 4 | Other | 2.1 |
| 4 | Other | 2 |
| 5 | rater1 | 2.5 |
| 5 | rater2 | 2.7 |
| 5 | Other | 2.1 |
期望数据结构
把第二个Other标签改为Other2:
| id | rater | raterscore |
|---|---|---|
| 1 | rater2 | 3 |
| 1 | rater1 | 2 |
| 1 | Other | 3 |
| 2 | rater1 | 1.5 |
| 2 | rater2 | 3.2 |
| 2 | Other | 1.1 |
| 2 | Other2 | 2 |
| 3 | rater1 | 2.5 |
| 3 | rater2 | 2.7 |
| 3 | Other | 2.1 |
| 3 | Other2 | 2 |
| 4 | rater1 | 2.5 |
| 4 | rater2 | 2.7 |
| 4 | Other | 2.1 |
| 4 | Other2 | 2 |
| 5 | rater1 | 2.5 |
| 5 | rater2 | 2.7 |
| 5 | Other | 2.1 |
解决方案
问题出在你先用filter(rater=="Other")筛选了数据,后续修改只针对这部分子集,没有合并回原数据。正确的做法是直接在原数据上分组处理,不用提前过滤:
方法一:用ifelse实现
library(dplyr) d_modified <- d %>% group_by(id, rater) %>% mutate( rater = ifelse(rater == "Other" & row_number() == 2, "Other2", rater) ) %>% ungroup()
方法二:用case_when实现(逻辑更清晰)
d_modified <- d %>% group_by(id, rater) %>% mutate( rater = case_when( rater == "Other" & row_number() == 2 ~ "Other2", TRUE ~ rater ) ) %>% ungroup()
核心逻辑说明
- 按
id和rater分组,这样每个id下的Other会单独成组,便于标记顺序 row_number()用来识别组内的行位置,第二个Other会被标记为2- 直接在原数据集上修改,无需过滤子集,确保所有原始数据都被保留,仅符合条件的rater标签被替换
内容的提问来源于stack exchange,提问作者helpneeder
相关产品推荐
相关产品推荐

