将数据框中因子Season的水平转为列(保留den原始值)
R数据框重塑:将Season因子水平转为对应den列的独立列
需求:将数据框中因子Season的两个水平(DRY、WET)转为两个独立列,列值直接取自den列,不做求和、均值等聚合计算。
尝试过的方法及问题
1. dplyr + tidyr 组合方法
执行代码:
df %>% group_by(Season) %>% mutate(index = row_number()) %>% pivot_wider(names_from = "Season", values_from = "den")
报错信息:
Error in `n()`: ! Must only be used inside data-masking verbs like `mutate()`, `filter()`, and `group_by()`. Run `rlang::last_trace()` to see where the error occurred.
2. spread函数方法
执行代码:
spread(df, Season, den)
报错信息:
Error in `spread()`: ! Each row of output must be identified by a unique combination of keys.
3. reshape函数方法
执行代码:
reshape(data=df, timevar="Season",idvar="assem",direction="wide" )
输出结果及警告:
assem den.DRY den.WET 1 Goldspotted Killifish 0.0000000 0.3333333 2 Far 0.6666667 0.3333333 3 Pal 0.0000000 1.3333333 4 Gulf Pipefish 0.0000000 0.0000000 18 Rainwater Killifish NA 1.0000000 Warning messages: 1: In reshapeWide(data, idvar = idvar, timevar = timevar, varying = varying, : multiple rows match for Season=DRY: first taken 2: In reshapeWide(data, idvar = idvar, timevar = timevar, varying = varying, : multiple rows match for Season=WET: first taken
期望输出格式
assem WET DRY Goldspotted Killifish NA 0.000 Goldspotted Killifish NA 0.000 Goldspotted Killifish NA 0.000 Goldspotted Killifish NA 0.000 Goldspotted Killifish NA 1.333 Goldspotted Killifish 0.333 NA Goldspotted Killifish 2 NA Goldspotted Killifish 0.333 NA
原始数据结构
structure(list(Season = structure(c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L), levels = c("DRY", "WET"), class = "factor"), assem = c("Goldspotted Killifish", "Far", "Pal", "Gulf Pipefish", "Pal", "Goldspotted Killifish", "Far", "Goldspotted Killifish", "Goldspotted Killifish", "Goldspotted Killifish", "Far", "Goldspotted Killifish", "Pal", "Gulf Pipefish", "Far", "Goldspotted Killifish", "Pal", "Rainwater Killifish", "Rainwater Killifish", "Goldspotted Killifish"), den = c(0, 0.666666666666667, 0, 0, 0, 0, 0, 0, 1.33333333333333, 0, 0.333333333333333, 0.333333333333333, 1.33333333333333, 0, 0, 2, 2.33333333333333, 1, 1.33333333333333, 0.333333333333333)), row.names = c(NA, -20L), class = "data.frame")
解决方案
方法1:直接条件赋值(简洁高效)
通过mutate结合条件判断直接生成WET和DRY列,保留原始数据的每一行:
library(dplyr) df_processed <- df %>% mutate( DRY = ifelse(Season == "DRY", den, NA), WET = ifelse(Season == "WET", den, NA) ) %>% select(assem, WET, DRY)
方法2:pivot_wider + 组内行号(符合tidyverse风格)
给每个assem+Season组添加行号,确保pivot_wider有唯一的行标识,避免重复键报错:
library(dplyr) library(tidyr) df_processed <- df %>% group_by(assem, Season) %>% mutate(row_id = row_number()) %>% ungroup() %>% pivot_wider(names_from = Season, values_from = den) %>% select(-row_id)
两种方法都能得到符合期望的输出,其中方法1更直观,适合仅需拆分两列的场景;方法2扩展性更强,适合Season有更多水平的情况。
内容的提问来源于stack exchange,提问作者Nate
相关产品推荐
相关产品推荐

