如何合并两个结构一致的DataFrame并在标签冲突时优先保留'Frequent'取值?
Merge Two DataFrames with Priority for 'Frequent' User Status
Got it, let's solve this problem where you need to merge two DataFrames and prioritize the Frequent value whenever there's a conflict in the Frequent User column for the same User ID. Here are two straightforward ways to do this in R:
Step 1: First, Let's Set Up Your Example Data
First, let's recreate the DataFrames you provided to work with:
# Create df1 df1 <- data.frame( user_id = c( "000f2b1e-dde2-4227-9122-d197c674dea8", "001abd72-53f1-436a-a26f-e3187543afb6", "001c1e12-a8f7-4459-9b23-6b8bdc5d2625", "002d8272-8ee2-4e1a-a523-5ba0813c20f9", "0037abe8-7623-4ac7-9fbb-7398f505c1f6", "003911c6-c013-43b7-9996-e771dbe3ac83" ), frequent_user = c("Frequent", "Infrequent", "Infrequent", "Infrequent", "Infrequent", "Infrequent"), stringsAsFactors = FALSE ) # Create df2 df2 <- data.frame( user_id = c( "000f2b1e-dde2-4227-9122-d197c674dea8", "001abd72-53f1-436a-a26f-e3187543afb6", "001c1e12-a8f7-4459-9b23-6b8bdc5d2625", "002d8272-8ee2-4e1a-a523-5ba0813c20f9", "0037abe8-7623-4ac7-9fbb-7398f505c1f6", "003911c6-c013-43b7-9996-e771dbe3ac83" ), frequent_user = c("Infrequent", "Infrequent", "Infrequent", "Infrequent", "Infrequent", "Infrequent"), stringsAsFactors = FALSE )
Method 1: Using dplyr (Clean & Intuitive)
If you're comfortable with the dplyr package, this is the most readable approach:
library(dplyr) # Combine both DataFrames into one combined_df <- bind_rows(df1, df2) # Group by user_id, and prioritize "Frequent" if present final_df <- combined_df %>% group_by(user_id) %>% summarise( frequent_user = ifelse(any(frequent_user == "Frequent"), "Frequent", "Infrequent"), .groups = "drop" # Remove grouping after summarizing )
How This Works:
bind_rows()stacks the two DataFrames on top of each other.group_by(user_id)groups all rows by each unique user ID.summarise()checks if any row in the group hasfrequent_user = "Frequent":- If yes, we set the final value to
Frequent. - If no, we keep
Infrequent.
- If yes, we set the final value to
Method 2: Base R (No External Packages)
If you prefer not to use dplyr, you can achieve the same result with base R functions:
# Combine both DataFrames combined_df <- rbind(df1, df2) # Split by user_id, process each group, then recombine final_df <- do.call(rbind, lapply(split(combined_df, combined_df$user_id), function(user_group) { data.frame( user_id = unique(user_group$user_id), frequent_user = if(any(user_group$frequent_user == "Frequent")) "Frequent" else "Infrequent", stringsAsFactors = FALSE ) })) # Reset row names to clean up the output rownames(final_df) <- NULL
How This Works:
rbind()combines the two DataFrames (same asbind_rows()in base R).split()breaks the combined DataFrame into a list where each element is a group of rows for one user ID.lapply()processes each user group: checks ifFrequentexists, then creates a single-row DataFrame for that user.do.call(rbind, ...)recombines all the processed user groups back into a single DataFrame.
Verify the Result
For your example data, the final final_df will have the first user's frequent_user set to Frequent (since df1 had that value), and all others will stay Infrequent—exactly what you need!
内容的提问来源于stack exchange,提问作者Metel Stairs
相关产品推荐
相关产品推荐

