You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于标签匹配两个数据框列表并执行n/TotalArea列除法运算

问题:匹配数据框列表并执行列除法运算

示例数据

# Example columns
Label <- c("Blue_001_Series009", "Blue_001_Series009", "Blue_001_Series009", "Blue_001_Series009","Red_001_Series008", "Red_001_Series008","Red_001_Series008","Red_001_Series008","Blue_002_Series009", "Blue_002_Series009","Blue_002_Series009","Blue_002_Series009")
Pred <- c("Pear", "Orange", "Apple", "Peach", "Pear", "Orange", "Apple", "Peach", "Pear", "Orange", "Apple", "Peach")
n <- c(10, 223, 890, 34, 78, 902, 34, 211, 1007,209, 330, 446)

# Make example data frame
data <- data.frame(Label, Pred, n)

# Split dataframe into a list of dataframes
df <- split(data, f = data$Label)  

# Second dataframe example columns
Label1 <- c("Red_001_Series008","Blue_001_Series009", "Blue_002_Series009")
TotalArea <- c(1904, 578, 7092)

# Make dataframe
data1 <- data.frame(Label1, TotalArea)
# Split dataframe into a list of dataframes
df1 <- split(data1, f = data1$Label1)   

需求描述

两个列表中的数据框拥有相同标签,但顺序不一致,需要完成:

  • 基于标签将df与df1中的数据框一一匹配
  • 对每个匹配的标签组,将df对应数据框的n列除以df1对应数据框的TotalArea列

示例输入输出

  • df片段:
Label                  Pred   n
1 Blue_001_Series009   Pear  10
2 Blue_001_Series009 Orange 223
3 Blue_001_Series009  Apple 890
4 Blue_001_Series009  Peach  34
  • df1片段:
Label1      TotalArea
2   Blue_001_Series009       578
  • 期望结果:
Blue_001_Series009 Pear / Blue_001_Series009 TotalArea
10 / 578 = 0.0173

Blue_001_Series009 Orange / Blue_001_Series009 TotalArea
223 / 578 = 0.3858

实际场景中列表包含数百个数据框,需要支持大规模数据处理。


解决方案

方法1:基础R实现(高效无依赖)

利用列表名称作为匹配键(split后列表名称即为标签),直接索引匹配,适合大规模数据处理。

# 获取两边都存在的共同标签,避免索引错误
common_labels <- intersect(names(df), names(df1))

# 遍历共同标签,执行除法运算
result_list <- lapply(common_labels, function(label) {
  current_df <- df[[label]]
  # 提取对应标签的TotalArea值(单个值)
  total_area <- df1[[label]]$TotalArea
  # 计算比值并添加新列
  current_df$ratio <- current_df$n / total_area
  # 保留需要的列
  current_df[, c("Label", "Pred", "n", "ratio")]
})

# 给结果列表命名
names(result_list) <- common_labels

# 可选:合并所有结果为单个数据框
combined_result <- do.call(rbind, result_list)

方法2:tidyverse(purrr)实现(简洁易读)

如果习惯tidyverse语法,用map2结合标签排序实现:

library(tidyverse)

# 按共同标签顺序重新排列两个列表
df_ordered <- df[common_labels]
df1_ordered <- df1[common_labels]

# 映射执行运算
result_list <- map2(df_ordered, df1_ordered, function(df_chunk, df1_chunk) {
  df_chunk %>%
    mutate(ratio = n / df1_chunk$TotalArea) %>%
    select(Label, Pred, n, ratio)
})

names(result_list) <- common_labels

# 可选:合并为单个数据框
combined_result <- bind_rows(result_list)

格式化输出(匹配示例样式)

如果需要输出示例中的文本格式,可添加遍历格式化逻辑:

for (label in common_labels) {
  current_df <- df[[label]]
  total_area <- df1[[label]]$TotalArea
  
  for (i in 1:nrow(current_df)) {
    cat(paste0(current_df$Label[i], " ", current_df$Pred[i], " / ", label, " TotalArea\n"))
    cat(paste0(current_df$n[i], " / ", total_area, " = ", round(current_df$n[i]/total_area, 4), "\n\n"))
  }
}

关键说明

  • 核心逻辑是用列表名称作为匹配键,彻底规避顺序不一致的问题
  • 先取intersect(names(df), names(df1))确保只处理两边都存在的标签,避免报错
  • 两种方法均适合大规模数据:基础R的lapply执行效率高,tidyverse语法更简洁易维护

内容的提问来源于stack exchange,提问作者MM1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 20:40:34