使用tidyverse按列名命名规则将R宽表重塑为指定结构长表
tidyverse 实现宽表转指定长表结构方案
前置依赖
核心用到tidyr的长宽转换、字段拆分函数和dplyr的基础处理函数,先加载依赖包:
library(tidyverse)
示例数据
你提供的测试数据如下:
df <- structure(list(`1.tes1` = 8.00073085700465, `1.tes2` = 8.08008192865136, `1.tes3` = 7.67643993710322, `1.tes4` = 4.40797764861845, `1.tes5` = 8.07887886210789, `1.tes6` = 7.5133416960745, `2.tes1` = 8.85382519278079, `2.tes2` = 7.69705180134625, `2.tes3` = 7.23033538475091, `2.tes4` = 8.14366028991503, `2.tes5` = 8.00207221069391, `2.tes6` = 7.04604929055087, `3.tes1` = 5.56967515444227, `3.tes2` = 6.81971790904382, `3.tes3` = 7.69285459160427, `3.tes4` = 7.29436429730407, `3.tes5` = 7.39693058270568, `3.tes6` = 6.6688956545532, `4.tes1` = 7.02870405956919, `4.tes2` = 7.89704902680482, `4.tes3` = 7.207699266581, `4.tes4` = 8.07642509042209, `4.tes5` = 9.12013776731989, `4.tes6` = 8.73388960806046), row.names = c(NA, -1L), class = c("tbl_df", "tbl", "data.frame"))
实现代码
result <- df %>% # 宽表转长表,把所有列转换为列名-值对 pivot_longer(cols = everything(), names_to = "col_name", values_to = "Value") %>% # 按.拆分列名,得到分组ID和带重复编号的变量字段 separate(col_name, into = c("Group", "var_rep"), sep = "\\.") %>% mutate( # 提取纯字母部分为变量名 Variable = str_extract(var_rep, "[a-zA-Z]+"), # 提取数字部分为重复试验ID Replication = as.integer(str_extract(var_rep, "\\d+")), # 按需保留三位小数,匹配示例输出格式 Value = round(Value, 3), # 若需要Group为整数类型,可去掉下行注释 # Group = as.integer(Group) ) %>% # 按要求调整列顺序 select(Group, Variable, Replication, Value)
输出效果
运行后得到的result完全符合要求的结构,前10行示例如下:
# A tibble: 24 × 4 Group Variable Replication Value <chr> <chr> <int> <dbl> 1 1 tes 1 8.00 2 1 tes 2 8.08 3 1 tes 3 7.68 4 1 tes 4 4.41 5 1 tes 5 8.08 6 1 tes 6 7.51 7 2 tes 1 8.85 8 2 tes 2 7.70 9 2 tes 3 7.23 10 2 tes 4 8.14 # ℹ 14 more rows
内容的提问来源于stack exchange,提问作者Tianjian Qin
相关产品推荐
相关产品推荐

