在R中展开列向量:实现日志时间戳与对应任务行对齐
Got it, let's walk through how to fix this log data alignment and do the unstack operation you need. First, let's recap your data structure: you've got alternating rows where each task line is immediately followed by its corresponding timestamp line—we need to shift those timestamps up to match their tasks, then restructure the data as needed.
Step 1: Get Your Data into R
First, let's assume your log data is either in a text file or you've already loaded it as a character vector. For example, here's how to replicate your sample data:
# Sample log data matching your format log_data <- c( "V1 1 TASK [include_vars]", "2 Thursday 05 April 2018 20:21:52 -0500 (0:00:00.429) 0:00:00.429", "3 TASK [include_vars]", "4 Thursday 05 April 2018 20:21:53 -0500 (0:00:00.289) 0:00:00.718", "5 TASK [include_vars]", "6 Thursday 05 April 2018 20:21:53 -0500 (0:00:00.270) 0:00:00.988" ) # If reading from a file instead: # log_data <- readLines("your_log_file.txt")
Step 2: Align Tasks with Their Timestamps
Since tasks are on odd-indexed rows (1, 3, 5...) and timestamps on even-indexed rows (2,4,6...), we can split them into separate vectors then combine them into a single data frame:
# Split into task rows and timestamp rows task_rows <- log_data[seq(1, length(log_data), by = 2)] timestamp_rows <- log_data[seq(2, length(log_data), by = 2)] # Create aligned data frame, clean up task names (remove leading numbers) aligned_df <- data.frame( Task = gsub("^V?\\d+\\s+", "", task_rows), # Strip leading IDs like "V1 1" or "3" Full_Timestamp = timestamp_rows, stringsAsFactors = FALSE )
After this, you'll have a clean data frame where each row has a task and its matching timestamp.
Step 3: Split Timestamp into Useful Columns (Optional)
Your timestamps have multiple pieces of info (entry number, datetime, elapsed time, total elapsed time). Let's split those into separate columns using tidyr::separate:
library(tidyr) aligned_df <- aligned_df %>% separate(Full_Timestamp, into = c("Entry_Num", "Datetime", "Elapsed_Time", "Total_Elapsed"), sep = "\\s+(?=\\(\\d|\\d:)") # Smart split to keep datetime intact # Clean up elapsed time (remove parentheses) aligned_df$Elapsed_Time <- gsub("[()]", "", aligned_df$Elapsed_Time)
Step 4: Unstack the Column Vector
If you need to unstack the data (e.g., pivot tasks into columns with their timestamp values), use base R's unstack() or tidyr functions. For example, if you add a run identifier to group entries:
# Add a run ID to group tasks (adjust based on your actual data) aligned_df$Run_ID <- 1:nrow(aligned_df) # Unstack to have tasks as columns, datetime as values unstacked_df <- unstack(aligned_df, Datetime ~ Task)
If you need a more flexible pivot, tidyr::pivot_wider works great too:
library(dplyr) unstacked_df <- aligned_df %>% pivot_wider(names_from = Task, values_from = Datetime)
Final Result
You'll end up with a structured data frame where each task is paired with its correct timestamp, and unstacked into the column format you need.
内容的提问来源于stack exchange,提问作者Alex Dometrius

