You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中实现分组行转列:将行转为列并以对应值填充

解决大型数据框分组转宽表问题

你需要把长格式的大型数据框按相似行分组,将指定字段转译为多列,本质是长表转宽表的操作,下面用具体示例演示正确实现方式。

原始数据示例

假设你的原始数据是长格式(id为分组依据,attribute是要转成列的字段,value是对应列的值):

library(tidyverse)

# 构造示例数据
original_df <- data.frame(
  id = rep(c(1, 2, 3), each = 3),
  attribute = c("height", "weight", "age", "height", "weight", "age", "height", "weight", "age"),
  value = c(175, 70, 30, 180, 80, 35, 165, 60, 28)
)

常见错误尝试

如果你的尝试代码只是简单分组后保留原结构,比如这样:

# 错误示例:无法实现转列效果
original_df %>%
  group_by(id) %>%
  mutate(attribute = attribute)

这样只会保留原长表结构,无法将attribute的不同值转为独立列。

正确实现方法

使用tidyr::pivot_wider()函数,直接完成长表转宽表的核心操作:

# 核心代码
result_df <- original_df %>%
  pivot_wider(
    id_cols = id,  # 指定分组依据列
    names_from = attribute,  # 指定要转为新列名的字段
    values_from = value  # 指定新列对应的值来源
  )

最终结果数据框

运行后得到的目标格式数据框如下:

> result_df
# A tibble: 3 × 4
     id height weight   age
  <dbl>  <dbl>  <dbl> <dbl>
1     1    175     70    30
2     2    180     80    35
3     3    165     60    28

带聚合的场景扩展

如果同一分组+属性字段存在重复值,需要先做聚合(如取均值、求和),可以先添加分组聚合步骤:

# 构造含重复值的示例数据
original_df_with_duplicates <- data.frame(
  id = rep(1, 4),
  attribute = c("height", "height", "weight", "weight"),
  value = c(175, 176, 70, 71)
)

# 先聚合再转宽
result_df_agg <- original_df_with_duplicates %>%
  group_by(id, attribute) %>%
  summarize(avg_value = mean(value), .groups = "drop") %>%
  pivot_wider(
    id_cols = id,
    names_from = attribute,
    values_from = avg_value
  )

内容的提问来源于stack exchange,提问作者MisterCoder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 11:20:27