如何在tidyverse中实现指定格式长表转换,保留m变量重复展示
在tidyverse中实现保留重复变量的长格式转换
问题说明
我可以用tidyr::pivot_longer(DATA, m:y, values_to= "z", names_to= "variable")将DATA转换为长格式得到variable列,但需要实现:每个唯一id对应的m变量,既要在variable=="y"的行中保留原值,又要在variable=="m"的行中作为z列的值重复出现,同时生成sm和sy两个指示变量。以下是用reshape2包实现的方案,现需用tidyverse工具完成相同需求。
数据示例
DATA <- read.csv("https://stats.idre.ucla.edu/stat/data/ml_sim.csv")
数据预览:
## id x m y ## 1 1 1.5451 0.1068 0.568 ## 2 1 2.2753 2.1104 1.206 . . .
期望输出Desired_output预览:
## fid id x m variable z sm sy ## 1 1 1 1.5451 0.1068 y 0.5678 0 1 ## 2 2 1 2.2753 2.1104 y 1.2061 0 1 ## 3 3 1 0.7867 0.0389 y -0.2613 0 1 ## ... ## 801 1 1 1.5451 0.1068 m 0.1068 1 0 ## 802 2 1 2.2753 2.1104 m 2.1104 1 0 ## ...
reshape2实现方案
# 使用reshape2的解决方案 stacked <- reshape2::melt(DATA, id.vars = c("id", "x", "m"), measure.vars = c("y", "m"), value.name = "z") Desired_output <- within(stacked, { sy <- as.integer(variable == "y") sm <- as.integer(variable == "m") })
tidyverse实现方案
方法一:pivot_longer直接转换+生成指示变量
核心是将m同时作为保留的id变量和待转换的measure变量,再通过mutate生成指示列:
library(tidyverse) Desired_output_tidy <- DATA %>% pivot_longer( cols = c(y, m), # 指定要转换的列:y和m names_to = "variable", values_to = "z" ) %>% mutate( sy = as.integer(variable == "y"), sm = as.integer(variable == "m") ) %>% arrange(variable, id) # 按variable和id排序,匹配示例输出顺序
方法二:拆分构造后合并
如果需要更直观的逻辑,可以分别生成variable="y"和variable="m"的数据集再合并:
library(tidyverse) # 构造variable=y的行 y_rows <- DATA %>% mutate( variable = "y", z = y, sy = 1, sm = 0 ) %>% select(id, x, m, variable, z, sm, sy) # 构造variable=m的行 m_rows <- DATA %>% mutate( variable = "m", z = m, sy = 0, sm = 1 ) %>% select(id, x, m, variable, z, sm, sy) # 合并并排序 Desired_output_tidy <- bind_rows(y_rows, m_rows) %>% arrange(variable, id)
两种方法均可得到与reshape2方案一致的输出,方法一更简洁,符合tidyverse的管道式编程风格。
内容的提问来源于stack exchange,提问作者Simon Harmel
相关产品推荐
相关产品推荐

