R语言新手求助:基于因子级别及相同ID创建条件列f
Hey there! No worries at all—we all start somewhere with R, and it’s totally okay to ask basic questions. Let’s figure this out together.
Looking at your data and desired output, the rule for column f is clear: for each group of rows sharing the same id, if the group contains both a row where b=0 and c=1 and a row where b=1 and c=0, set f=1 for all rows in that group; otherwise, set f=0. And crucially, we’ll keep all your original columns (like e) intact.
Solution using dplyr (tidyverse)
First, install the dplyr package if you haven’t already with install.packages("dplyr"), then use this code:
library(dplyr) # Your sample data frame (I’ve structured it properly here) df <- tibble( b = c(0, 1, 0, 1), c = c(1, 0, 1, 0), id = c(45, 45, 48, 46), e = c(5, 7, 5, 7) ) # Add the new column f without altering existing columns df_with_f <- df %>% group_by(id) %>% mutate( # Check if both required (b,c) pairs exist in the group has_b0_c1 = any(b == 0 & c == 1), has_b1_c0 = any(b == 1 & c == 0), # Convert the combined condition to 1/0 for column f f = as.integer(has_b0_c1 & has_b1_c0) ) %>% # Remove helper columns (optional but keeps data clean) select(-has_b0_c1, -has_b1_c0) %>% ungroup() # View the final result df_with_f
How this works:
group_by(id): Clusters rows by theiridvalue so we can check conditions per groupmutate(): Creates new columns within each group:has_b0_c1andhas_b1_c0are helper flags to check if each required (b,c) pair exists in the groupfis set to 1 only if both flags areTRUE, converted to an integer for your desired 1/0 format
select(-...): Cleans up the helper columns we used to calculatefungroup(): Returns the data to a regular ungrouped data frame
Base R alternative (no packages needed)
If you prefer to stick to base R without installing extra packages, use this approach with ave():
# Your sample data frame df <- data.frame( b = c(0, 1, 0, 1), c = c(1, 0, 1, 0), id = c(45, 45, 48, 46), e = c(5, 7, 5, 7) ) # Add column f directly to the original data frame df$f <- with(df, ave( x = 1:nrow(df), by = id, FUN = function(group_rows) { # Subset the data to the current id group group_data <- df[group_rows, ] # Check if both (b,c) pairs exist, return 1 or 0 as.integer(any(group_data$b == 0 & group_data$c == 1) & any(group_data$b == 1 & group_data$c == 0)) } )) # View the result df
Both methods will preserve all your original columns while adding the f column exactly as you need it. Let me know if you hit any snags getting this to work!
内容的提问来源于stack exchange,提问作者prospectjoe

