You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何根据数据框列最大值为列分配Cluster分组?求dplyr函数建议

基于列最大值分配样本到聚类分组的R实现方案

需求说明

根据数据框中各样本列(Sample1、Sample2、Sample3)的最大值,将每列分配到对应的Cluster分组(Cluster1、Cluster2、Cluster3)中,优先推荐dplyr相关工具实现。

示例数据

原始数据结构如下:

Group                Sample1         Sample2        Sample3
1 Cluster1             0.1             0               0.1
2 Cluster2             0.4             0.3             0.01
3 Cluster3             0               0.9             0.92

对应的R数据框结构代码:

df <- structure(list(Group = c("Cluster1", "Cluster2", "Cluster3"), 
  Sample1 = c(0.1, 0.4, 0), Sample2 = c(0, 0.3, 0.9), Sample3 = c(0.1, 
  0.01, 0.92)), class = "data.frame", row.names = c("1", "2", "3"))

预期输出

Sample1 Cluster2
Sample2 Cluster3
Sample3 Cluster3

实现方案

方法1:使用dplyr + tidyr(推荐)

通过数据格式转换+分组筛选实现,代码如下:

library(dplyr)
library(tidyr)

# 处理数据并生成结果
result <- df %>%
  # 将宽格式转为长格式,便于按样本分组处理
  pivot_longer(-Group, names_to = "Sample", values_to = "Value") %>%
  # 按每个样本列分组
  group_by(Sample) %>%
  # 筛选出每个样本组中值最大的行
  slice_max(Value, n = 1) %>%
  # 取消分组
  ungroup() %>%
  # 保留需要的两列
  select(Sample, Group)

# 打印结果(隐藏行名)
print(result, row.names = FALSE)

方法2:Base R 实现

无需额外包,直接通过矩阵操作完成:

# 计算每个样本列最大值所在的行索引
max_row_indices <- apply(df[, -1], 2, which.max)

# 组合成结果数据框
result_base <- data.frame(
  Sample = colnames(df[, -1]),
  Group = df$Group[max_row_indices]
)

# 打印结果(隐藏行名)
print(result_base, row.names = FALSE)

内容的提问来源于stack exchange,提问作者user2300940

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 08:40:15