使用plyr包dlply/ldply后,如何按Subject ID生成起始为1的序列列
Hey there! Let's walk through this with a concrete example so you can see exactly how to add that incrementing column you need, while keeping your existing workflow with plyr's dlply and ldply.
First, let's start with a reproducible dataset to mimic your scenario:
library(plyr) # Create sample data matching your structure set.seed(123) df <- data.frame( `Subject IDs` = rep(c(10, 11, 12), each = 3), time_seq = rep(c(1, 2, 3), 3), measurement = rnorm(9) # Random values to represent your data )
Step 1: Add the incrementing column first (simplest approach)
Before splitting your data, you can generate the group-specific incrementing column using ddply to group by Subject IDs and add a sequence number for each row in the group:
# Add a new column "group_seq" that starts at 1 for each Subject ID df_with_seq <- ddply(df, .(`Subject IDs`), transform, group_seq = seq_along(`Subject IDs`))
The seq_along(Subject IDs) part will generate a sequence from 1 to the number of rows in each Subject group—perfect for your "start at 1 and increment" requirement.
Step 2: Split into named lists (as you originally did)
Now split the modified data frame by Subject IDs and time_seq to get your desired named list (like 10.1, 10.2):
split_list <- dlply(df_with_seq, .(`Subject IDs`, time_seq), identity) # Check the list names to confirm they match your format names(split_list) # Output: "10.1" "10.2" "10.3" "11.1" "11.2" "11.3" "12.1" "12.2" "12.3"
Step 3: Merge back to a data frame with ldply
Finally, use ldply to combine the list back into a single data frame, which will include your new group_seq column:
final_df <- ldply(split_list, identity) # View the result head(final_df)
This will give you output like this:
Subject IDs time_seq measurement group_seq 1 10 1 -0.5604756 1 2 10 2 -0.2301775 2 3 10 3 1.5587083 3 4 11 1 0.0705084 1 5 11 2 0.1292877 2 6 11 3 1.7150650 3
If you need to add the column after splitting
If you must add the column post-split (e.g., if your list is already generated), you can adjust the workflow to first split by Subject IDs, add the sequence to each group, then split further by time_seq:
# Split by Subject IDs first split_by_subject <- dlply(df, .(`Subject IDs`), function(subject_data) { # Add the incrementing column to each subject's data subject_data$group_seq <- seq_len(nrow(subject_data)) subject_data }) # Split each subject's data by time_seq to get the ID.time names final_split_list <- ldply(split_by_subject, function(x) dlply(x, .(time_seq), identity)) %>% unlist(recursive = FALSE) # Fix the names to match ID.time format names(final_split_list) <- sapply(strsplit(names(final_split_list), "\\."), function(parts) paste(parts[2], parts[1], sep = ".")) # Merge back to data frame final_df <- ldply(final_split_list, identity)
Either way, you'll end up with the incrementing column grouped by Subject IDs that you need.
内容的提问来源于stack exchange,提问作者Frederick

