基于含唯一值的数据框填充含重复项的数据框
Got it, let's work through this together. First, let's recap your unique ID dataframe setup to make sure we're on the same page:
set.seed(1) df_Unique <- matrix(rnorm(20),4,5) colnames(df_Unique) <- paste0("k_",1:5) rownames(df_Unique) <- paste0("Project",1:4) # 转换为标准数据框方便后续操作 df_Unique <- as.data.frame(df_Unique)
First, we need to adjust the unique dataframe to include project IDs as a proper column (instead of just rownames) — this makes matching with your duplicate dataframe much smoother. Then we can use a join operation to fill in the missing values.
Step 1: Simulate your duplicate dataframe
Let's assume your duplicate dataframe looks like this (it has repeated Project IDs and needs the k_1 to k_5 values filled in):
df_Duplicates <- data.frame(ProjectID = c("Project1", "Project2", "Project1", "Project3", "Project4", "Project2"))
Step 2: Option 1 — Base R Solution
We'll use merge() to perform a left join, which preserves all rows (including duplicates) from your duplicate dataframe and pulls in matching data from the unique dataframe:
# 把行名转为单独的ProjectID列 df_Unique_with_id <- df_Unique df_Unique_with_id$ProjectID <- rownames(df_Unique_with_id) # 左连接填充:保留df_Duplicates的所有行,匹配对应数据 filled_df <- merge(df_Duplicates, df_Unique_with_id, by = "ProjectID", all.x = TRUE)
Step 3: Option 2 — dplyr Solution (More Intuitive)
If you prefer tidyverse syntax, left_join() from dplyr works perfectly here and is easier to read:
library(dplyr) # 转换行名为列,并把ProjectID移到第一列方便查看 df_Unique_tidy <- df_Unique %>% mutate(ProjectID = rownames(.)) %>% relocate(ProjectID) # 执行左连接,自动根据ProjectID匹配填充 filled_df <- df_Duplicates %>% left_join(df_Unique_tidy, by = "ProjectID")
What You'll Get
The resulting filled_df will keep all rows from your duplicate dataframe, with each repeated Project ID paired with the corresponding k_1 to k_5 values from df_Unique. For example, the two "Project1" rows will both have the values -0.6264538, 0.3295078, 0.5757814, -0.62124058, -0.01619026 in the k columns.
内容的提问来源于stack exchange,提问作者Sam

