You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中去除重复行并将信息保留至另一列?

方法一:使用tidyverse工具链(dplyr + tidyr)

这是最直观简便的实现方式,通过将长格式数据转为宽格式来合并重复id的信息:

library(tidyverse)

df <- data.frame(id = c("a", "a", "b", "c", "c", "d"),
                 color = c("red", "blue", "green", "blue", "green","red"))

df2 <- df %>%
  group_by(id) %>%
  mutate(color_col = paste0("color", row_number())) %>%
  pivot_wider(names_from = color_col, values_from = color) %>%
  rename(color = color1) %>%
  replace(is.na(.), "")

df2

运行后即可得到目标结果,各步骤作用:

  • group_by(id):按id分组处理重复项
  • mutate:为每个分组内的color生成带序号的列名标识
  • pivot_wider:将长格式数据转换为宽格式,把同一id的多个color拆分到不同列
  • rename:把生成的color1重命名为color,匹配需求格式
  • replace:将空值(NA)替换为空白字符串

方法二:Base R 原生实现

如果不想加载额外扩展包,可使用基础R的函数完成:

df <- data.frame(id = c("a", "a", "b", "c", "c", "d"),
                 color = c("red", "blue", "green", "blue", "green","red"))

# 为每个id内的行生成序号
df$time <- ave(df$id, df$id, FUN = seq_along)
# 转换为宽格式
df2 <- reshape(df, idvar = "id", timevar = "time", direction = "wide")
# 清理列名
names(df2) <- gsub("color\\.", "", names(df2))
# 将空值替换为空白字符串
df2[is.na(df2)] <- ""

df2

内容的提问来源于stack exchange,提问作者Victor Shin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 00:53:14