You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:修改分组后重复的最后一行指定列值

需求说明

我有一个DataFrame,第一列为日期,包含数月数据。已通过复制每个月最后一行的方式生成了重复行,现在需要修改这些重复行的指定列值:将每月最后一行的副本(即重复行)的file列值更新为下一个分组的file值,使其成为下一组的首行。

原数据片段

# A tibble: 11 × 3
   date        file value
   <date>     <int> <dbl>
 1 2021-07-26     1   858
 2 2021-07-27     1   880
 3 2021-07-28     1   904
 4 2021-07-29     1   933
 5 2021-07-30     1   970
 6 2021-07-31     1  1010
 7 2021-07-31     1  1010
 8 2021-08-01     2  1048
 9 2021-08-02     2  1078
10 2021-08-03     2  1096
11 2021-08-04     2  1107

预期输出

# A tibble: 11 × 3
   date        file value
   <date>     <int> <dbl>
 1 2021-07-26     1   858
 2 2021-07-27     1   880
 3 2021-07-28     1   904
 4 2021-07-29     1   933
 5 2021-07-30     1   970
 6 2021-07-31     1  1010
 7 2021-07-31     2  1010  # file changed from 1 to 2 as desired
 8 2021-08-01     2  1048
 9 2021-08-02     2  1078
10 2021-08-03     2  1096
11 2021-08-04     2  1107

数据结构

df <- structure(list(date = structure(c(18834, 18835, 18836, 18837, 
18838, 18839, 18839, 18840, 18841, 18842, 18843), class = "Date"), 
    file = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L), value = c(858, 
    880, 904, 933, 970, 1010, 1010, 1048, 1078, 1096, 1107)), row.names = c(NA, 
-11L), class = c("tbl_df", "tbl", "data.frame"))
解决方案

方式一:基于重复行位置与固定递增规律

如果file值是按月份固定递增+1,直接标记重复行并更新:

library(dplyr)

df_modified <- df %>%
  group_by(date) %>%
  # 标记同一日期的第二行(即重复生成的行)
  mutate(is_dup = row_number() == 2) %>%
  ungroup() %>%
  # 对重复行的file值加1
  mutate(file = ifelse(is_dup, file + 1, file)) %>%
  select(-is_dup)

print(df_modified)

方式二:匹配后续分组的实际file值(通用场景)

如果file的递增规律不固定,需要匹配下一组的实际file值,可通过关联分组起始日期实现:

library(dplyr)

# 提取每个file对应的最早日期,生成下一个file的映射
file_map <- df %>%
  distinct(file, .keep_all = TRUE) %>%
  arrange(date) %>%
  mutate(next_file = lead(file))

df_modified <- df %>%
  group_by(date) %>%
  mutate(is_dup = row_number() == 2) %>%
  ungroup() %>%
  left_join(file_map %>% select(file, next_file), by = "file") %>%
  mutate(file = ifelse(is_dup, next_file, file)) %>%
  select(-is_dup, -next_file)

print(df_modified)

内容的提问来源于stack exchange,提问作者wernor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 06:54:23