You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中合并含相同列的数据集并对相同列值求和

R数据集合并:保留唯一列+相同列求和+替换NaN为0

需求说明

合并mem和birds两个数据集:

  • 保留两个数据集的所有唯一列
  • 对两数据集共有的列(如3z)按行求和
  • 所有NaN值替换为0

示例数据集

先构建可复现的示例数据:

# 构建birds数据集
birds <- data.frame(
  `1a` = c(10, 2, 4, 2, 0),
  `2c` = c(0, 0, 0, 2, 0),
  `3z` = c(10, 2, 13, 1, 0),
  `4x` = c(10, 4, 0, 0, 0)
)
rownames(birds) <- 1:5

# 构建mem数据集
mem <- data.frame(
  `13a` = c(8, 0, 3, 2, 1),
  `7x` = c(7, 0, 0, 3, 0),
  `3z` = c(11, 2, 1, 1, 1),
  `6z` = c(1, 7, 8, 0, 0)
)
rownames(mem) <- 6:10

解决方案

方法1:基础R实现(无需额外包)

# 获取所有唯一列名
all_columns <- unique(c(colnames(birds), colnames(mem)))

# 扩展两个数据集至所有列,NA填充为0
birds_expanded <- birds[, all_columns]
birds_expanded[is.na(birds_expanded)] <- 0

mem_expanded <- mem[, all_columns]
mem_expanded[is.na(mem_expanded)] <- 0

# 合并并按行名排序
final_output <- rbind(birds_expanded, mem_expanded)
final_output <- final_output[order(as.integer(rownames(final_output))), ]

print(final_output)

方法2:dplyr/tidyr实现(适合复杂场景)

先加载依赖包:

library(dplyr)
library(tidyr)

然后执行合并操作:

# 转成长格式并合并
combined_data <- bind_rows(
  birds %>% rownames_to_column("row_id") %>% pivot_longer(-row_id),
  mem %>% rownames_to_column("row_id") %>% pivot_longer(-row_id)
) %>%
  group_by(row_id, name) %>%
  summarise(value = sum(value, na.rm = TRUE), .groups = "drop") %>%
  pivot_wider(names_from = name, values_from = value) %>%
  mutate(across(everything(), ~replace_na(.x, 0))) %>%
  arrange(row_id)

# 设置行名并移除辅助列
rownames(combined_data) <- combined_data$row_id
combined_data$row_id <- NULL

print(combined_data)

两种方法都能得到你需要的期望输出,运行后即可看到结果。

内容的提问来源于stack exchange,提问作者pradoer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 18:53:20