根据最后出现情况修正数据框中机器操作序列
高效修正机器操作序列的R方案
原数据
df <- data.frame( Machine_ID= c(1,1,1,1,2,2,2,2,2,2,3,3,3), Operation = c("Start", "Stop", "Start", "Stop", "Stop", "Stop", "Start", "Stop", "Start", "Stop", "Start", "Stop", "Start") )
需求说明
- 每台机器的操作次数为偶数
- 严格遵循
Start与Stop成对的序列规则 - 修正开头错误的
Stop,补充末尾缺失的Stop
高效实现代码(基于dplyr)
无需循环,利用分组向量化操作处理,适合大数据集:
library(dplyr) corrected_df <- df %>% group_by(Machine_ID) %>% summarise( Operation = { ops <- Operation # 修正开头的错误Stop:将首项Stop替换为Start,保证第一对有效 if (length(ops) >= 1 && ops[1] == "Stop") { ops[1] <- "Start" } # 将序列拆分为两两一组,确保每组都是Start/Stop,奇数长度则补Stop paired_groups <- split(ops, ceiling(seq_along(ops) / 2)) fixed_pairs <- lapply(paired_groups, function(group) { if (length(group) == 1) { c(group, "Stop") } else { c("Start", "Stop") } }) unlist(fixed_pairs) }, .groups = "drop" ) %>% # 重置行号匹配期望格式 mutate(` ` = row_number()) %>% select(` `, Machine_ID, Operation) # 查看结果 print(corrected_df, row.names = FALSE)
输出结果
Machine_ID Operation 1 1 Start 2 1 Stop 3 1 Start 4 1 Stop 5 2 Start 6 2 Stop 7 2 Start 8 2 Stop 9 2 Start 10 2 Stop 11 3 Start 12 3 Stop 13 3 Start 14 3 Stop
方案优势
- 避免循环操作,利用dplyr的分组优化逻辑,处理大规模数据时性能远高于循环
- 逻辑清晰,通过分组拆分+组内修正,确保每台机器的序列严格符合成对规则
- 自动处理奇数长度的序列,补充缺失的
Stop
内容的提问来源于stack exchange,提问作者onhalu
相关产品推荐
相关产品推荐

