在R语言中重塑多列数据框:宽表转长表
Got it, let's tackle this data reshaping problem! You've got a wide dataframe where every 3 columns represent a sample, and you want to convert it into a long format with sample, wave, cdom, and total columns. Here are a few straightforward ways to do this using different R tools:
Using Tidyverse (dplyr + tidyr)
This approach uses regex to parse column names efficiently, making it easy to group columns into samples:
library(tidyverse) # Your original dataframe df <- data.frame( wave = c(350, 352), cdom = c(0.164910664183534, 0.161336423973549), total = c(0.173292853508359, 0.164541380243188), wave_1 = c(350, 352), cdom_1 = c(0.157738707282744, 0.149555740098184), total_1 = c(0.16501632769282, 0.151631889636391), wave_2 = c(350, 352), cdom_2 = c(0.143293704793142, 0.133057094683334), total_2 = c(0.148878497119496, 0.136150629840465), wave_3 = c(350, 352), cdom_3 = c(0.0972284241775975, 0.0906890150335725), total_3 = c(0.108645612944463, 0.103640164204995), wave_4 = c(350, 352), cdom_4 = c(0.0801780489449968, 0.0779336395415438), total_4 = c(0.103930690374372, 0.095768602460239) ) # Reshape to long format df_long <- df %>% pivot_longer( cols = everything(), names_to = c(".value", "sample"), names_pattern = "(wave|cdom|total)(_\\d+)?" ) %>% # Fix the first sample (no _ suffix in original columns) mutate(sample = ifelse(is.na(sample), "1", str_remove(sample, "_"))) %>% # Convert sample to numeric for cleaner sorting mutate(sample = as.numeric(sample)) %>% arrange(sample, wave) # View the result df_long
Output:
#> sample wave cdom total #> 1 1 350 0.1649107 0.1732929 #> 2 1 352 0.1613364 0.1645414 #> 3 2 350 0.1577387 0.1650163 #> 4 2 352 0.1495557 0.1516319 #> 5 3 350 0.1432937 0.1488785 #> 6 3 352 0.1330571 0.1361506 #> 7 4 350 0.0972284 0.1086456 #> 8 4 352 0.0906890 0.1036402 #> 9 5 350 0.0801780 0.1039307 #> 10 5 352 0.0779336 0.0957686
Using Base R
If you prefer avoiding external packages, the built-in reshape() function works well:
# Reshape with base R df_long_base <- reshape( df, direction = "long", # Group columns by their prefixes varying = list( wave = grep("wave", names(df)), cdom = grep("cdom", names(df)), total = grep("total", names(df)) ), # Name the resulting value columns v.names = c("wave", "cdom", "total"), # Define the sample identifier column timevar = "sample", # Assign sample numbers 1 to 5 times = 1:5, # Create an id for original rows to avoid conflicts idvar = "original_row", new.row.names = 1:(nrow(df)*5) ) %>% # Clean up columns and sort select(sample, wave, cdom, total) %>% arrange(sample, wave) df_long_base
This produces the same clean long-format output as the tidyverse method.
Using data.table
For fast reshaping (ideal for large datasets), data.table's melt() function is a efficient choice:
library(data.table) # Convert to data.table setDT(df) # Reshape to long format df_long_dt <- melt( df, # Match columns by their prefixes measure.vars = patterns("^wave", "^cdom", "^total"), # Name the value columns value.name = c("wave", "cdom", "total"), # Name the sample identifier column variable.name = "sample" ) %>% arrange(sample, wave) df_long_dt
Again, this gives the same desired long-format dataframe.
内容的提问来源于stack exchange,提问作者Philippe Massicotte

