如何基于start_time与time_end_of_experiment为每个参与者生成end_time变量(R语言)
Create
end_time Variable for Each Participant in R Perfect, let's tackle this problem. You want to create an end_time column where each entry is either:
- The next row's
start_timeminus 10ms (0.01 seconds) for all but the last row per participant, or - The
time_end_of_experimentminus 10ms for the final row of each subject.
Here are two reliable approaches to achieve your desired output:
Method 1: Using dplyr (Tidyverse Style)
This method is intuitive and clean, especially for grouped operations:
# Load the dplyr package (install first if needed: install.packages("dplyr")) library(dplyr) # Your initial data frame df <- data.frame( subject_nr = c("1", "1", "1", "2", "2"), start_time = c(50, 52, 55, 53, 54.5), time_end_of_experiment = c(60, 60, 60, 55.5, 55.5) ) # Create the end_time column df <- df %>% group_by(subject_nr) %>% # Group data by each participant mutate( end_time = ifelse( row_number() == n(), # Check if this is the last row of the group time_end_of_experiment - 0.01, # Use experiment end time minus 10ms lead(start_time) - 0.01 # Use next row's start time minus 10ms ) ) %>% ungroup() # Remove grouping to keep data frame in standard format # View the result print(df)
Explanation:
group_by(subject_nr): Ensures we process each participant's data separately.row_number() == n(): Identifies the final row in each participant's group.lead(start_time): Fetches thestart_timevalue from the next row in the same group.ungroup(): Resets the data frame to an ungrouped state (good practice for subsequent operations).
Method 2: Using Base R
If you prefer not to use external packages, this base R approach works just as well:
# Your initial data frame df <- data.frame( subject_nr = c("1", "1", "1", "2", "2"), start_time = c(50, 52, 55, 53, 54.5), time_end_of_experiment = c(60, 60, 60, 55.5, 55.5) ) # Calculate end_time using tapply to handle grouped logic df$end_time <- unlist(tapply( seq_len(nrow(df)), # Use row indices to group data df$subject_nr, function(idx) { sub_df <- df[idx, ] # Get next start time (NA for last row of the group) next_start <- c(sub_df$start_time[-1], NA) # Assign end_time values ifelse( is.na(next_start), sub_df$time_end_of_experiment - 0.01, next_start - 0.01 ) } )) # View the result print(df)
Explanation:
tapply(): Groups row indices bysubject_nrand applies a custom function to each group.next_start <- c(sub_df$start_time[-1], NA): Creates a vector where each entry is the next row'sstart_time, with NA for the last row.ifelse(): Replaces NA values withtime_end_of_experiment - 0.01and usesnext_start - 0.01for all other rows.
Final Output
Both methods will produce your desired data frame:
subject_nr start_time time_end_of_experiment end_time 1 1 50.0 60.0 51.9 2 1 52.0 60.0 54.9 3 1 55.0 60.0 59.9 4 2 53.0 55.5 54.4 5 2 54.5 55.5 55.4
内容的提问来源于stack exchange,提问作者user9974638
相关产品推荐
相关产品推荐

