R语言:修改分组后重复的最后一行指定列值
需求说明
我有一个DataFrame,第一列为日期,包含数月数据。已通过复制每个月最后一行的方式生成了重复行,现在需要修改这些重复行的指定列值:将每月最后一行的副本(即重复行)的file列值更新为下一个分组的file值,使其成为下一组的首行。
原数据片段
# A tibble: 11 × 3 date file value <date> <int> <dbl> 1 2021-07-26 1 858 2 2021-07-27 1 880 3 2021-07-28 1 904 4 2021-07-29 1 933 5 2021-07-30 1 970 6 2021-07-31 1 1010 7 2021-07-31 1 1010 8 2021-08-01 2 1048 9 2021-08-02 2 1078 10 2021-08-03 2 1096 11 2021-08-04 2 1107
预期输出
# A tibble: 11 × 3 date file value <date> <int> <dbl> 1 2021-07-26 1 858 2 2021-07-27 1 880 3 2021-07-28 1 904 4 2021-07-29 1 933 5 2021-07-30 1 970 6 2021-07-31 1 1010 7 2021-07-31 2 1010 # file changed from 1 to 2 as desired 8 2021-08-01 2 1048 9 2021-08-02 2 1078 10 2021-08-03 2 1096 11 2021-08-04 2 1107
数据结构
df <- structure(list(date = structure(c(18834, 18835, 18836, 18837, 18838, 18839, 18839, 18840, 18841, 18842, 18843), class = "Date"), file = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L), value = c(858, 880, 904, 933, 970, 1010, 1010, 1048, 1078, 1096, 1107)), row.names = c(NA, -11L), class = c("tbl_df", "tbl", "data.frame"))
解决方案
方式一:基于重复行位置与固定递增规律
如果file值是按月份固定递增+1,直接标记重复行并更新:
library(dplyr) df_modified <- df %>% group_by(date) %>% # 标记同一日期的第二行(即重复生成的行) mutate(is_dup = row_number() == 2) %>% ungroup() %>% # 对重复行的file值加1 mutate(file = ifelse(is_dup, file + 1, file)) %>% select(-is_dup) print(df_modified)
方式二:匹配后续分组的实际file值(通用场景)
如果file的递增规律不固定,需要匹配下一组的实际file值,可通过关联分组起始日期实现:
library(dplyr) # 提取每个file对应的最早日期,生成下一个file的映射 file_map <- df %>% distinct(file, .keep_all = TRUE) %>% arrange(date) %>% mutate(next_file = lead(file)) df_modified <- df %>% group_by(date) %>% mutate(is_dup = row_number() == 2) %>% ungroup() %>% left_join(file_map %>% select(file, next_file), by = "file") %>% mutate(file = ifelse(is_dup, next_file, file)) %>% select(-is_dup, -next_file) print(df_modified)
内容的提问来源于stack exchange,提问作者wernor
相关产品推荐
相关产品推荐

