如何在R语言中按批次替换dataframe列的‘shifting cycle’字符串
问题描述
现有DataFrame某列内容如下:
first cycle first cycle shifting cycle shifting cycle shifting cycle 2nd cycle 2nd cycle 2nd cycle shifting cycle shifting cycle 3rd cycle 3rd cycle
需将第一组连续的shifting cycle替换为shifting cycle 1,第二组替换为shifting cycle 2。目前通过关联其他列手动赋值的方式处理,但多文件数据存在差异,该方法不通用。
现有代码:
df$column <- str_replace(df$column, "Shifting cycle", "Shifting cycle 2") df <- df %>% mutate(column = case_when(other_column ==30~ 'Shifting cycle 1' ,T~column))
预期输出:
first cycle first cycle shifting cycle 1 shifting cycle 1 shifting cycle 1 2nd cycle 2nd cycle 2nd cycle shifting cycle 2 shifting cycle 2 3rd cycle 3rd cycle
解决方案
可以利用rle()函数识别连续重复的分组,为每组shifting cycle分配递增编号后替换字符串,无需依赖其他列,适配多文件场景。
具体代码实现
library(dplyr) library(stringr) # 示例数据(实际使用时替换为你的DataFrame) df <- tibble( column = c( "first cycle", "first cycle", "shifting cycle", "shifting cycle", "shifting cycle", "2nd cycle", "2nd cycle", "2nd cycle", "shifting cycle", "shifting cycle", "3rd cycle", "3rd cycle" ) ) # 生成连续组的统计信息 rle_result <- rle(df$column) # 标记shifting cycle的组并分配编号 is_shifting <- rle_result$values == "shifting cycle" shift_group_ids <- cumsum(is_shifting) * is_shifting # 将编号映射到每一行并替换字符串 df <- df %>% mutate( column = if_else( column == "shifting cycle", str_c(column, rep(shift_group_ids, rle_result$lengths), sep = " "), column ) )
代码说明
rle()函数会返回目标列中连续重复值的长度和对应值,帮我们识别出所有连续的shifting cycle分组;- 通过
cumsum(is_shifting)为每个shifting cycle分组生成递增的编号(1、2、3...); - 用
rep()将分组编号扩展为与原数据行数一致的向量,再通过str_c()拼接字符串完成替换。
内容的提问来源于stack exchange,提问作者شاه نواز
相关产品推荐
相关产品推荐

