R数据重塑:将Nest行值转为列字段以适配Wilcox检验需求
数据转换实现方案
你需要完成的是长表转宽表操作,以下是可直接运行的R代码实现:
1. tidyverse 实现(推荐,代码可读性更高)
library(tidyverse) # 你的样本数据 df <- structure(list(BreedingSeason = c("2017-2018", "2018-2019", "2016-2017", "2014-2015", "2017-2018", "2017-2018"), Nest = c(0L, 0L, 1L, 1L, 1L, 1L), HubDist = c(828.4829,3125.2952, 1433.3238, 1539.5785, 1602.5289, 1431.7102)), row.names = c(NA, 6L), class = "data.frame") # 转换为目标格式 df_wide <- df %>% # 按繁殖季、Nest分组,为同组内的多条数据分配唯一序号,避免转宽时聚合丢失数据 group_by(BreedingSeason, Nest) %>% mutate(row_id = row_number()) %>% ungroup() %>% # 长转宽,Nest取值作为列名,对应值为HubDist pivot_wider( id_cols = c(BreedingSeason, row_id), names_from = Nest, values_from = HubDist ) %>% # 重命名繁殖季列,删除辅助序号列 rename(BreedingS = BreedingSeason) %>% select(-row_id)
转换后的数据结构如下:
| BreedingS | 0 | 1 |
|---|---|---|
| 2017-2018 | 828.4829 | 1433.324 |
| 2017-2018 | NA | 1602.529 |
| 2017-2018 | NA | 1431.710 |
| 2018-2019 | 3125.295 | NA |
| 2016-2017 | NA | 1433.324 |
| 2014-2015 | NA | 1539.579 |
2. 纯base R实现(无需加载第三方包)
# 你的样本数据 df <- structure(list(BreedingSeason = c("2017-2018", "2018-2019", "2016-2017", "2014-2015", "2017-2018", "2017-2018"), Nest = c(0L, 0L, 1L, 1L, 1L, 1L), HubDist = c(828.4829,3125.2952, 1433.3238, 1539.5785, 1602.5289, 1431.7102)), row.names = c(NA, 6L), class = "data.frame") # 生成辅助序号列 df$row_id <- ave(df$HubDist, df$BreedingSeason, df$Nest, FUN = seq_along) # 长转宽 df_wide <- reshape(df, idvar = c("BreedingSeason", "row_id"), timevar = "Nest", direction = "wide") # 调整列名 names(df_wide) <- gsub("HubDist\\.", "", names(df_wide)) names(df_wide)[1] <- "BreedingS" df_wide$row_id <- NULL
后续使用说明
- 做Wilcox检验时,直接提取
df_wide$"0"和df_wide$"1"列,去掉NA值即可执行检验 - 需要按繁殖季筛选时,直接对应过滤
BreedingS列的取值即可
内容的提问来源于stack exchange,提问作者Burton Guster
相关产品推荐
相关产品推荐

