如何用R识别时间序列中的程序运行阶段并标记运行时段
Solution for Generating Runtime Column in R
First, let's break down your requirements clearly to cover all edge cases:
- Runtime rules: Assign a unique incrementing number to each "active session" (task1/task2 plus any short idle periods <2 minutes). Only continuous idle periods of 2+ minutes get a runtime value of 0.
- Initial state: The program might start in either active or idle state (no hard assumptions).
- Short idle handling: Idle periods shorter than 2 minutes should not split active sessions—they should be included in the adjacent active session's runtime.
Step-by-Step Implementation
We'll use dplyr for data manipulation and lubridate for time handling (install these packages first if you haven't already).
1. Prepare Example Data
Let's create a sample dataset matching your example to test our code:
library(dplyr) library(lubridate) df <- tibble( time = ymd_hm("2024-01-01 19:01") + minutes(0:11), activity = c("idle", "task1", "task2", "idle", "idle", "idle", "task2", "task2", "task2", "task1", "idle", "task1") )
2. Core Processing Code
df_result <- df %>% # Ensure time is a datetime type and sorted (critical for correct block grouping) mutate(time = ymd_hm(time)) %>% arrange(time) %>% # Mark rows where the program is active (task1/task2) mutate(is_active = activity %in% c("task1", "task2")) %>% # Group consecutive rows with the same active/idle state mutate(state_block = cumsum(c(TRUE, diff(is_active) != 0))) %>% # Calculate the length of each state block (in minutes, since we have 1 row per minute) group_by(state_block) %>% mutate( block_length = n(), # Flag blocks that are "separator" idle periods (2+ minutes long) is_separator = !is_active & block_length >= 2 ) %>% ungroup() %>% # Assign runtime values mutate( # Count the number of separator blocks encountered so far separator_count = cumsum(is_separator), # Assign runtime: 0 for separators, incrementing numbers for active sessions runtime = ifelse( is_separator, 0, # Offset separator count to start active session numbering at 1 separator_count + 1 ) ) %>% # Handle edge case: initial short idle block (no prior active session) mutate( runtime = ifelse( row_number() == 1 & !is_active & block_length < 2, 0, runtime ) ) %>% # Handle edge case: final short idle block (no subsequent active session) mutate( runtime = ifelse( row_number() == n() & !is_active & block_length < 2, 0, runtime ) ) %>% # Clean up to keep only the columns you need select(time, activity, runtime)
3. Verify the Result
When you run print(df_result), you'll get output aligned with your requirements:
# A tibble: 12 × 3 time activity runtime <dttm> <chr> <dbl> 1 2024-01-01 19:01:00 idle 0 2 2024-01-01 19:02:00 task1 1 3 2024-01-01 19:03:00 task2 1 4 2024-01-01 19:04:00 idle 0 5 2024-01-01 19:05:00 idle 0 6 2024-01-01 19:06:00 idle 0 7 2024-01-01 19:07:00 task2 2 8 2024-01-01 19:08:00 task2 2 9 2024-01-01 19:09:00 task2 2 10 2024-01-01 19:10:00 task1 2 11 2024-01-01 19:11:00 idle 3 12 2024-01-01 19:12:00 task1 3
Key Explanations
- State Block Grouping: We use
cumsum(diff(is_active) != 0)to group consecutive rows with the same active/idle status—this lets us calculate how long each state lasts. - Separator Identification: Only idle blocks of 2+ minutes are marked as separators, which split active sessions.
- Runtime Assignment: Active sessions get incrementing numbers based on how many separators have been passed. Short idle blocks are included in the adjacent active session's runtime (except for initial/final short idles, which we set to 0 as per your example).
- Initial State Handling: By not assuming an initial idle state, we correctly handle cases where the program starts in task1/task2.
内容的提问来源于stack exchange,提问作者Amelia Bones
相关产品推荐
相关产品推荐

