R语言中识别允许容错间隔的连续序列技术咨询
rle() Hey Joel! Great job getting the basic rle() workflow down for spotting runs of 1s in your data. Let's expand that to handle those "forgivable" NA gaps you mentioned—where a small number of NAs between 1s should count as part of the same continuous event sequence.
Quick Recap of Your Original Code
First, let's confirm your existing setup works as expected:
temp.data = rle(c(NA, NA, 1, NA, NA, 1, NA, 1, 1, 1, NA, NA, NA)) output = temp.data$lengths[temp.data$value==1] # Output: [1] 1 1 3
This correctly picks out the strict runs of 1s, ignoring any NAs as full breaks. Now let's add the fault tolerance.
Solution: Allow Controlled NA Gaps
The core idea is to first "fill in" NAs that are within your allowed gap size between 1s, then run rle() as before. Here's a flexible implementation where you set the maximum number of NAs you'll tolerate between 1s:
# Your raw data (example with mixed gaps) raw_data <- c(NA, NA, 1, NA, 1, NA, NA, 1, 1, 1, NA, 1) # Set your maximum allowed NA gap (adjust this to your needs!) max_allowed_gap <- 1 # Step 1: Mark all positions that are 1 or valid gap NAs is_event_related <- rep(FALSE, length(raw_data)) is_event_related[raw_data == 1] <- TRUE # Step 2: Check each NA to see if it's within the allowed gap of 1s for (i in which(is.na(raw_data))) { # Find nearest 1 to the left and right nearest_left <- ifelse(any(is_event_related[1:i]), max(which(is_event_related[1:i])), 0) nearest_right <- ifelse(any(is_event_related[i:length(raw_data)]), min(which(is_event_related[i:length(raw_data)])) + i - 1, length(raw_data) + 1) # If the NA is close enough to both sides (or one side if at the edge), mark it as part of the event if ((i - nearest_left <= max_allowed_gap + 1) && (nearest_right - i <= max_allowed_gap + 1)) { is_event_related[i] <- TRUE } } # Step 3: Convert valid positions to 1, others to NA processed_data <- ifelse(is_event_related, 1, NA) # Step 4: Run rle on the processed data temp_data <- rle(processed_data) output <- temp_data$lengths[temp_data$value == 1] # Check the output output # For the example raw_data, output is [1] 3 3 1 # Explanation: The first 1, NA, 1 becomes a run of 3; the 1,1,1,NA,1 becomes a run of 3; the final 1 is a run of 1
How It Works
max_allowed_gap: This is your control knob—set it to 0 for strict detection (matches your original code), 1 for allowing single NA gaps, 2 for two NAs, etc.- The loop checks each NA to see if it's within your gap limit of a 1 on either side. If yes, it treats that NA as part of the event sequence.
- After processing, we use the familiar
rle()function to get the lengths of the now "fault-tolerant" event runs.
Bonus: Simplified Alternative for Small Gaps
If you only need to tolerate 1 NA gap, you can use a simpler rolling window approach with zoo::rollapply() (though you'll need to install the zoo package first):
library(zoo) raw_data <- c(NA, NA, 1, NA, 1, NA, NA, 1, 1, 1, NA, 1) processed_data <- ifelse(rollapply(raw_data, 3, function(x) any(x == 1), fill = FALSE), 1, NA) temp_data <- rle(processed_data) output <- temp_data$lengths[temp_data$value == 1]
This marks any position (including NAs) that has a 1 in its immediate 3-position window as part of the event sequence.
内容的提问来源于stack exchange,提问作者Joel

