You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并两个结构一致的DataFrame并在标签冲突时优先保留'Frequent'取值?

Merge Two DataFrames with Priority for 'Frequent' User Status

Got it, let's solve this problem where you need to merge two DataFrames and prioritize the Frequent value whenever there's a conflict in the Frequent User column for the same User ID. Here are two straightforward ways to do this in R:

Step 1: First, Let's Set Up Your Example Data

First, let's recreate the DataFrames you provided to work with:

# Create df1
df1 <- data.frame(
  user_id = c(
    "000f2b1e-dde2-4227-9122-d197c674dea8", 
    "001abd72-53f1-436a-a26f-e3187543afb6", 
    "001c1e12-a8f7-4459-9b23-6b8bdc5d2625", 
    "002d8272-8ee2-4e1a-a523-5ba0813c20f9", 
    "0037abe8-7623-4ac7-9fbb-7398f505c1f6", 
    "003911c6-c013-43b7-9996-e771dbe3ac83"
  ),
  frequent_user = c("Frequent", "Infrequent", "Infrequent", "Infrequent", "Infrequent", "Infrequent"),
  stringsAsFactors = FALSE
)

# Create df2
df2 <- data.frame(
  user_id = c(
    "000f2b1e-dde2-4227-9122-d197c674dea8", 
    "001abd72-53f1-436a-a26f-e3187543afb6", 
    "001c1e12-a8f7-4459-9b23-6b8bdc5d2625", 
    "002d8272-8ee2-4e1a-a523-5ba0813c20f9", 
    "0037abe8-7623-4ac7-9fbb-7398f505c1f6", 
    "003911c6-c013-43b7-9996-e771dbe3ac83"
  ),
  frequent_user = c("Infrequent", "Infrequent", "Infrequent", "Infrequent", "Infrequent", "Infrequent"),
  stringsAsFactors = FALSE
)

Method 1: Using dplyr (Clean & Intuitive)

If you're comfortable with the dplyr package, this is the most readable approach:

library(dplyr)

# Combine both DataFrames into one
combined_df <- bind_rows(df1, df2)

# Group by user_id, and prioritize "Frequent" if present
final_df <- combined_df %>%
  group_by(user_id) %>%
  summarise(
    frequent_user = ifelse(any(frequent_user == "Frequent"), "Frequent", "Infrequent"),
    .groups = "drop"  # Remove grouping after summarizing
  )

How This Works:

  • bind_rows() stacks the two DataFrames on top of each other.
  • group_by(user_id) groups all rows by each unique user ID.
  • summarise() checks if any row in the group has frequent_user = "Frequent":
    • If yes, we set the final value to Frequent.
    • If no, we keep Infrequent.

Method 2: Base R (No External Packages)

If you prefer not to use dplyr, you can achieve the same result with base R functions:

# Combine both DataFrames
combined_df <- rbind(df1, df2)

# Split by user_id, process each group, then recombine
final_df <- do.call(rbind, lapply(split(combined_df, combined_df$user_id), function(user_group) {
  data.frame(
    user_id = unique(user_group$user_id),
    frequent_user = if(any(user_group$frequent_user == "Frequent")) "Frequent" else "Infrequent",
    stringsAsFactors = FALSE
  )
}))

# Reset row names to clean up the output
rownames(final_df) <- NULL

How This Works:

  • rbind() combines the two DataFrames (same as bind_rows() in base R).
  • split() breaks the combined DataFrame into a list where each element is a group of rows for one user ID.
  • lapply() processes each user group: checks if Frequent exists, then creates a single-row DataFrame for that user.
  • do.call(rbind, ...) recombines all the processed user groups back into a single DataFrame.

Verify the Result

For your example data, the final final_df will have the first user's frequent_user set to Frequent (since df1 had that value), and all others will stay Infrequent—exactly what you need!

内容的提问来源于stack exchange,提问作者Metel Stairs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 02:47:48