在R语言中重塑数据,将count列转换为多列格式
R语言:将分类列转换为多列并填充对应值
问题描述
现有如下data.frame数据:
df<-structure(list(concept = c("agree", "anger", "anxiety", "cognitive", "cognitive"), count = c(1L, 2L, 4L, 6L, 122L)), class = "data.frame", row.names = c(NA, -5L))
需要为concept列的每个唯一取值创建新列,对应行的count值填入新列,其余行填充0,最终目标格式如下:
new_df<-structure(list(agree = c(1L, 0L, 0L, 0L, 0L), anger = c(0L, 2L, 0L, 0L, 0L), anxiety = c(0L, 0L, 4L, 0L, 0L), cognitive = c(0L, 0L, 0L, 6L, 122L)), class = "data.frame", row.names = c(NA, -5L ))
输出预览:
agree anger anxiety cognitive 1 1 0 0 0 2 0 2 0 0 3 0 0 4 0 4 0 0 0 6 5 0 0 0 122
解决方案
方法1:Base R 原生实现
循环填充法
# 获取所有唯一的分类值 unique_concepts <- unique(df$concept) # 初始化全0数据框 new_df <- data.frame(matrix(0, nrow = nrow(df), ncol = length(unique_concepts))) colnames(new_df) <- unique_concepts # 遍历每个分类,填充对应值 for (concept in unique_concepts) { new_df[df$concept == concept, concept] <- df$count[df$concept == concept] }
model.matrix 简洁法
# 添加行ID确保顺序不变 df$row_id <- seq_len(nrow(df)) # 生成分类矩阵并乘以count值 new_df <- model.matrix(~ concept - 1, df) * df$count # 合并行ID并恢复原顺序 new_df <- cbind(new_df, df$row_id) colnames(new_df)[ncol(new_df)] <- "row_id" new_df <- new_df[order(new_df$row_id), ] # 清理行ID和行名 rownames(new_df) <- NULL new_df$row_id <- NULL
方法2:tidyverse 工具包实现
使用dplyr和tidyr的pivot_wider函数快速转换:
library(tidyverse) new_df <- df %>% # 添加行ID保证转换后顺序一致 mutate(row_id = row_number()) %>% pivot_wider( id_cols = row_id, names_from = concept, # 作为新列名的列 values_from = count, # 填充值的列 values_fill = 0 # 缺失值填充为0 ) %>% select(-row_id) # 移除行ID列
方法3:data.table 工具包实现
用dcast函数高效完成转换:
library(data.table) # 转换为data.table格式 setDT(df) # 宽表转换,.I代表原行索引保证顺序,fill=0填充空值 new_df <- dcast(df, .I ~ concept, value.var = "count", fill = 0) # 移除行索引列 new_df[, .I := NULL]
内容的提问来源于stack exchange,提问作者Ali Roghani
相关产品推荐
相关产品推荐

