使用R统计数字1的连续出现批次及单次持续时长的技术问询
Hey there! Let's work through your two problems step by step using R and dplyr (since you already used pipe operators in your attempts).
Your goal is to assign a unique number to each block of consecutive 1s in the var column, and set count to 0 where var is 0. Your initial rle approach was on the right track, but we need to expand it to map the run-length encoding back to each row in your dataset.
Solution Code
library(dplyr) # Your sample data dt <- tibble( time = c(0, 0.3, 0.6, 0.9, 1.2, 1.5, 1.8), var = c(1, 1, 0, 0, 1, 1, 0) ) # Generate the count column dt <- dt %>% # Create a group ID for each consecutive block of the same var value mutate(group_id = cumsum(var != lag(var, default = var[1]))) %>% # Calculate the length of each group and its var value group_by(group_id) %>% mutate( group_length = n(), group_var = first(var) ) %>% ungroup() %>% # Assign unique batch IDs to var=1 groups; set 0 for var=0 mutate( count = ifelse( group_var == 1, dense_rank(ifelse(group_var == 1, group_id, NA)), 0 ) ) %>% # Clean up temporary columns (optional) select(-group_id, -group_length, -group_var) print(dt)
Output
# A tibble: 7 × 3 time var count <dbl> <dbl> <dbl> 1 0 1 1 2 0.3 1 1 3 0.6 0 0 4 0.9 0 0 5 1.2 1 2 6 1.5 1 2 7 1.8 0 0
Notes
If you only want to assign IDs to blocks of 3 or more consecutive 1s (matching your original description of "连续次数多于2次"), modify the count line to:
count = ifelse( group_var == 1 & group_length > 2, dense_rank(ifelse(group_var == 1 & group_length > 2, group_id, NA)), 0 )
Your previous code had two main issues: using a string default for lag() (mismatched with numeric var values) and grouping by time (which creates a group for every single row). Here's the fix:
Solution Code
# Add the replicate column (per your note) dt$replicate <- 1 # Calculate batch durations batch_duration <- dt %>% # Reuse the group ID logic from Part 1 to identify consecutive blocks mutate(group_id = cumsum(var != lag(var, default = var[1]))) %>% # Group by replicate and the continuous block ID group_by(replicate, group_id) %>% # Keep only blocks where var is 1 filter(var == 1) %>% # Calculate start/end times and duration summarise( start_time = min(time), end_time = max(time), duration = end_time - start_time, .groups = "drop" ) %>% # Optional: Add batch IDs to match Part 1 mutate(batch_id = row_number()) print(batch_duration)
Output
# A tibble: 2 × 5 replicate group_id start_time end_time duration batch_id <dbl> <int> <dbl> <dbl> <dbl> <int> 1 1 1 0 0.3 0.3 1 2 1 3 1.2 1.5 0.3 2
Explanation
- We use the same
group_idlogic to group consecutive rows with the samevarvalue. - Grouping by
replicateensures we handle multiple replicates correctly (even though your current data only hasreplicate=1). summarise()calculates the start, end, and duration for each valid 1-block.
内容的提问来源于stack exchange,提问作者hamjam223

