在R中基于均值与置信区间使用pivot_longer实现指定列转长
问题
我需要生成包含var、attend、mean、ci_lower、ci_upper这5列的结果数据框来绘制系数图,但当前使用pivot_longer会把整个数据框转成长格式,如何仅针对attend相关列执行转长操作?
原代码:
dtest_long <- dtest %>% pivot_longer(!var, names_to = "attend", values_to = "value")
数据:
dtest <- structure(list(var = c("jbj_ext_1", "jbj_ext_2", "jbj_ext_3", "jbj_int_1", "jbj_int_2", "jbj_int_3", "jbj_tot_1", "jbj_tot_2", "jbj_tot_3"), no_attend_mean = c(11.15, 8, 7.64, 4.24, 8.26, 7.89, 30.01, 21.08, 20.93), attend_mean = c(15.36, 10.9, 11.15, 10.1, 4.25, 4.24, 28.98, 20.99, 19.89), no_attend_lowCI = c(9.65, 6.5, 6.14, 2.74, 6.76, 6.39, 28.51, 19.58, 19.43), attend_lowCI = c(13.86, 9.4, 9.65, 8.6, 2.75, 2.74, 27.48, 19.49, 18.39), no_attend_upperCI = c(12.2, 9.05, 8.69, 5.29, 9.31, 8.94, 31.06, 22.13, 21.98), attend_upperCI = c(16.41, 11.95, 12.2, 11.15, 5.3, 5.29, 30.03, 22.04, 20.94)), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -9L))
解决方案
可以通过pivot_longer的正则匹配拆分列名,再结合pivot_wide整理成目标格式:
library(tidyr) library(dplyr) dtest_target <- dtest %>% # 拆分列名为attend分组和指标类型 pivot_longer( cols = -var, names_to = c("attend", "metric"), # 正则表达式匹配:前缀(no_attend/attend) + 下划线 + 后缀(mean/lowCI/upperCI) names_pattern = "(.*)_(mean|lowCI|upperCI)", values_to = "value" ) %>% # 将指标类型转成列 pivot_wider( names_from = metric, values_from = value ) %>% # 重命名CI列到目标名称 rename( ci_lower = lowCI, ci_upper = upperCI )
执行后得到的数据框结构如下(以前3行为例):
# A tibble: 18 × 5 var attend mean ci_lower ci_upper <chr> <chr> <dbl> <dbl> <dbl> 1 jbj_ext_1 no_attend 11.1 9.65 12.2 2 jbj_ext_1 attend 15.4 13.9 16.4 3 jbj_ext_2 no_attend 8 6.5 9.05
这样就得到了包含var、attend、mean、ci_lower、ci_upper5列的目标数据框,正好满足绘制系数图的需求。
内容的提问来源于stack exchange,提问作者a_todd12
相关产品推荐
相关产品推荐

