R语言如何在for循环的sample()操作中添加约束保证time序列连续
你只需要调整filter判断的时间范围即可实现需求,改动后的完整代码如下:
library(tidyverse) data <- expand_grid(study=1:4,group=1:2, time=0:4) n = 1 time_to_remove <- unique(data$time)[-(1:2)] unique_study <- unique(data$study) for(i in time_to_remove) { studies_to_remove = sample(unique_study, size = n) time_to_remove = i unique_study <- setdiff(unique_study, studies_to_remove) data <- data %>% filter( !(study %in% studies_to_remove & time >= time_to_remove) # 仅改动这一处,把原逻辑的`time %in% time_to_remove`改为`time >= time_to_remove` ) } data %>% as.data.frame()
改动原理
原来的逻辑是仅删除选中study的time=i数据,会导致被删time=2的study后续还保留time=3、4,出现序列断裂。
改为time >= time_to_remove后,被选中的study会直接删除所有大于等于当前待删时间点的所有数据,自动保证剩余的time都是从0开始的连续前缀:
- 被选中删
time=2的study:仅保留0、1 - 被选中删
time=3的study:仅保留0、1、2 - 被选中删
time=4的study:仅保留0、1、2、3 - 全程没被选中的study:保留0-4全部时间点
完全符合你要求的4种合法剩余序列,不会出现跳断情况。
内容的提问来源于stack exchange,提问作者Simon Harmel
相关产品推荐
相关产品推荐

