如何在Count值匹配时用df2更新df1中患者的IntDen值
问题:根据Count值匹配更新IntDen列数据
需求说明
我有两个数据集df1和df2,需要针对每位患者实现:
- 对比df1中Count行的各个value列值,与df2中同一患者Count行的所有value列值
- 若df1某value列的Count值在df2的Count行中存在,则将df1对应患者IntDen行的该value列值,替换为df2中同一患者IntDen行、匹配到的那个value列的值
尝试的代码
我写了以下代码,但无法得到正确的更新结果:
# Create sample data frames df1 <- data.frame(patient = c(rep("A", 5), rep("B", 5), rep("C", 5)), measurement = rep(c("Count", "Area", "StdDev", "IntDen", "RawIntDen"), 3), value1 = 1:15, value2 = 2:16, value3 = 3:17, value4 = 4:18, value5 = 5:19) df2 <- data.frame(patient = c(rep("A", 5), rep("B", 5), rep("C", 5)), measurement = rep(c("Count", "Area", "StdDev", "IntDen", "RawIntDen"), 3), value1 = c(rep(1,15)), value2 = c(rep(4, 15)), value3 = 5:19, value4 = 6:20, value5 = 7:21, value6 = 8:22) # Create a copy of df1 to store the updated values df1_updated <- df1 # Loop over each patient in df1 for (patient in unique(df1$patient)) { # Subset df1 and df2 for the current patient df1_patient <- df1[df1$patient == patient, ] df2_patient <- df2[df2$patient == patient, ] # Get the row index of the "Count" measurement in df1 and df2 count_row_idx_df1 <- which(df1_patient$measurement == "Count") count_row_idx_df2 <- which(df2_patient$measurement == "Count") # Get the row index of the "IntDen" measurement in df1 and df2 intden_row_idx_df1 <- which(df1_patient$measurement == "IntDen") intden_row_idx_df2 <- which(df2_patient$measurement == "IntDen") # Check if the "Count" value in df1 is in df2 count_value <- df1_patient[count_row_idx_df1, grep("value", names(df1_patient))] for(count_value in df2_patient[count_row_idx_df2, grep("value", names(df2_patient))]) { if (any(count_value) %in% df2_patient[count_row_idx_df2, grep("value", names(df2_patient))]) { # Get the column index of the matching "Count" value in df2 match_col_idx_df2 <- which(df2_patient[count_row_idx_df2, grep("value", names(df2_patient))] %in% count_value) # Update the "IntDen" value in df1_updated intden_value <- df2_patient[intden_row_idx_df2, match_col_idx_df2] df1_updated[intden_row_idx_df1, match_col_idx_df2] <- intden_value } } }
这段代码会修改一些值,但结果都是错误的。
期望的输出
正确的df1_updated可通过以下代码生成:
df1_updated_example <- data.frame(patient = c(rep("A", 5), rep("B", 5), rep("C", 5)), measurement = rep(c("Count", "Area", "StdDev", "IntDen", "RawIntDen"), 3), value1 = c(1,2,3,1,5:15), value2 = 2:16, value3 = 3:17, value4 = c(4,5,6,4,8:18), value5 = 5:19)
输出解释
仅患者A的IntDen行有修改:
- df1中患者A的Count行value1值为1,该值在df2患者A的Count行value1中存在,因此将df1患者A的IntDen行value1值从4改为1
- df1中患者A的Count行value4值为4,该值在df2患者A的Count行value2中存在,因此将df1患者A的IntDen行value4值从6改为4
解决方案
你的代码问题在于循环逻辑混乱,错误地遍历了df2的Count值,且没有正确对应df1的列和df2的匹配列。以下是修正后的代码:
# 创建示例数据 df1 <- data.frame(patient = c(rep("A", 5), rep("B", 5), rep("C", 5)), measurement = rep(c("Count", "Area", "StdDev", "IntDen", "RawIntDen"), 3), value1 = 1:15, value2 = 2:16, value3 = 3:17, value4 = 4:18, value5 = 5:19) df2 <- data.frame(patient = c(rep("A", 5), rep("B", 5), rep("C", 5)), measurement = rep(c("Count", "Area", "StdDev", "IntDen", "RawIntDen"), 3), value1 = c(rep(1,15)), value2 = c(rep(4, 15)), value3 = 5:19, value4 = 6:20, value5 = 7:21, value6 = 8:22) df1_updated <- df1 # 遍历每个患者 for (p in unique(df1$patient)) { # 提取当前患者的df1和df2数据 df1_p <- df1[df1$patient == p, ] df2_p <- df2[df2$patient == p, ] # 获取Count和IntDen行的位置 df1_count_row <- which(df1_p$measurement == "Count") df1_intden_row <- which(df1_p$measurement == "IntDen") df2_count_row <- which(df2_p$measurement == "Count") df2_intden_row <- which(df2_p$measurement == "IntDen") # 获取df1的value列和对应Count值 df1_value_cols <- grep("value", names(df1_p)) df1_count_vals <- df1_p[df1_count_row, df1_value_cols] # 获取df2的value列和Count值 df2_value_cols <- grep("value", names(df2_p)) df2_count_vals <- df2_p[df2_count_row, df2_value_cols] # 遍历df1的每个value列 for (col_idx in seq_along(df1_value_cols)) { current_val <- df1_count_vals[col_idx] # 查找df2中Count值匹配的列 match_cols <- which(df2_count_vals == current_val) if (length(match_cols) > 0) { # 取第一个匹配的列对应的IntDen值(若多个匹配可根据需求调整) new_intden_val <- df2_p[df2_intden_row, df2_value_cols[match_cols[1]]] # 更新df1_updated对应位置 # 找到当前患者在df1_updated中的行范围 patient_rows <- which(df1_updated$patient == p) df1_updated[patient_rows[df1_intden_row], df1_value_cols[col_idx]] <- new_intden_val } } } # 查看结果 df1_updated
代码说明
- 明确遍历df1的每个value列,确保对应关系正确
- 找到当前患者在
df1_updated中的完整行范围,避免子集行索引导致的定位错误 - 匹配到df2的Count值后,取对应列的IntDen值更新到df1的对应位置
内容的提问来源于stack exchange,提问作者Schulthe
相关产品推荐
相关产品推荐

