在R中重塑成对数据:按disease拆分e列为两列
R数据重塑:按disease拆分e列为两列
可以通过以下几种方法实现需求:
方法1:使用tidyr包(tidyverse系列)
这是现代R数据处理的常用直观方式:
首先加载数据:
df <- structure(list(id = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10), fid = c(1, 1, 2, 2, 3, 3, 4, 4, 5, 5), disease = c(0, 1, 0, 1, 1, 0, 1, 0, 0, 1), e = c(3, 2, 6, 1, 2, 5, 2, 3, 1, 1)), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -10L))
安装并加载tidyr包,然后用pivot_wider()完成重塑:
# 仅首次使用需要安装 # install.packages("tidyr") library(tidyr) reshaped_df <- df %>% pivot_wider( id_cols = c(id, fid), # 保留的标识列 names_from = disease, # 用于生成新列名的列 values_from = e, # 填充新列的数值来源 names_prefix = "e_" # 给新列名加前缀,避免纯数字列名 )
生成的reshaped_df中,e_0对应disease=0的e值,e_1对应disease=1的e值。
方法2:Base R原生函数
无需额外安装包,用reshape()函数实现:
reshaped_df_base <- reshape( df, idvar = c("id", "fid"), timevar = "disease", direction = "wide" )
该方法生成的列名为e.0和e.1,核心效果与方法1一致,仅列名格式略有差异。
输出示例
运行代码后打印结果:
print(reshaped_df)
输出结果如下:
# A tibble: 10 × 4 id fid e_0 e_1 <dbl> <dbl> <dbl> <dbl> 1 1 1 3 2 2 2 1 NA NA 3 3 2 6 1 4 4 2 NA NA 5 5 3 NA 2 6 6 3 5 NA 7 7 4 NA 2 8 8 4 3 NA 9 9 5 1 NA 10 10 5 NA 1
内容的提问来源于stack exchange,提问作者Lisa
相关产品推荐
相关产品推荐

