如何在R中将(n*P)行N列的data.frame转换为n行(N*P)列?
Problem Description
I'm working with a data.frame df in R that has n*P rows and N columns. The structure looks like this (with row names indicating group-position pairs x-y, where x is the group number (1 to n) and y is the position within the group (1 to P)):
# Example structure (simulated) set.seed(123) n <- 3 P <- 4 N <- 4 df <- data.frame( C1 = sample(1:200, n*P), C2 = sample(-50:100, n*P), CN_minus_1 = sample(50:200, n*P), CN = sample(10:100, n*P), row.names = paste0(rep(1:n, each = P), "-", rep(1:P, n)) )
I need to reshape this into a new data.frame df.new with n rows and N*P columns, following these rules:
- Each row
Rkindf.newis formed by column-binding (cbind) all rows from groupkindf(i.e., rowsk-1,k-2, ...,k-P). - Each block of N columns in
df.newcorresponds to the same position across all groups: the first N columns are rows1-1,2-1, ...,n-1row-bound (rbind), the next N columns are rows1-2,2-2, ...,n-2row-bound, and so on until the last N columns.
The target structure should look like this:
# Target structure (conceptual) # R1: values from df rows 1-1, 1-2, 1-3, 1-P (cbind into one row) # R2: values from df rows 2-1, 2-2, 2-3, 2-P (cbind into one row) # Rn: values from df rows n-1, n-2, n-3, n-P (cbind into one row) # Columns: C1-1, C2-1, ..., CN-1, C1-2, C2-2, ..., CN-2, ..., C1-P, C2-P, ..., CN-P
I tried using nested for loops but it didn't work well, here's my attempt:
for (j in 1:n) { df.new <- data.frame(matrix(vector(), 1, dim(df)[2], dimnames = list(c(), colnames(df))), stringsAsFactors=F) for (i in 1:nrow(df)) { if (i %% n == 0) { df.new <- rbind(df.new, df[i,]) } else if (i %% n == j) { df.new <- rbind(df.new, df[i,]) } } assign(paste0("df.new", j), df.new) }
Solution
Your loop approach is inefficient because repeated rbind operations copy the entire data frame each time, which gets slow with large datasets. Instead, use vectorized operations or tidyverse tools for a cleaner, faster solution.
Method 1: Base R (No External Packages)
This approach leverages matrix transposition and list operations to reshape the data efficiently:
# Step 1: Split the original data into n groups (each with P rows) group <- rep(1:n, each = P) df_groups <- split(df, group) # Step 2: Transpose each group to turn P rows into 1 row, then combine all groups df.new <- do.call(rbind, lapply(df_groups, function(g) { as.data.frame(t(g), stringsAsFactors = FALSE) })) # Step 3: Set meaningful column and row names colnames(df.new) <- paste0(rep(colnames(df), P), "-", rep(1:P, each = N)) rownames(df.new) <- paste0("R", 1:n)
Method 2: Tidyverse (dplyr + tidyr)
If you prefer a more readable, pipe-based workflow, use the tidyverse packages:
library(dplyr) library(tidyr) # Step 1: Add group and position columns to the original data df_with_groups <- df %>% mutate( group = rep(1:n, each = P), pos = rep(1:P, n) ) # Step 2: Reshape from wide to long, then back to wide in the desired format df.new <- df_with_groups %>% pivot_longer(cols = starts_with("C"), names_to = "column", values_to = "value") %>% pivot_wider( id_cols = group, names_from = c(column, pos), values_from = value, names_sep = "-" ) %>% column_to_rownames("group") # Step 3: Set row names (matches your target structure) rownames(df.new) <- paste0("R", 1:n)
Both methods will produce the exact structure you need, with better performance and maintainability than nested loops.
内容的提问来源于stack exchange,提问作者Makoto Miyazaki

