如何将多组x、y列的DataFrame转换为单组x、y列的DataFrame?
Alright, let's solve this DataFrame reshaping task using tidyr::gather and tidyr::spread as you asked—no loops needed! Here's a step-by-step breakdown with example code you can follow directly.
Step 1: Prepare Example Data (for testing)
First, let's create a sample a DataFrame that matches your column structure, so you can replicate the process:
library(tidyr) library(dplyr) # For helper functions like `select` and pipe operator # Sample DataFrame a <- data.frame( x1 = c(1, 2, 3), y1 = c(4, 5, 6), x2 = c(7, 8, 9), y2 = c(10, 11, 12), x3 = c(13, 14, 15), y3 = c(16, 17, 18) )
Step 2: Reshape Step-by-Step
1. Convert Wide to Long Format with gather
First, we use gather to collapse all columns into two: one for the column names (e.g., x1, y1) and one for their corresponding values:
a_long <- gather(a, key = "variable", value = "value")
This gives us a long DataFrame where each row represents a single value from the original columns, paired with its column name.
2. Split Column Names into Type and Group
Next, we split the variable column into two parts: the type (x or y) and the group number (1, 2, 3). We use tidyr::separate for this:
a_separated <- separate(a_long, col = variable, into = c("type", "group"), sep = 1)
The sep = 1 argument tells R to split the string after the first character (since our column names start with x/y followed by a number).
3. Reshape Back to Wide Format (x & y Columns)
Now we use spread to turn the type column into two separate columns (x and y), using the value column for their data:
a_reshaped <- spread(a_separated, key = type, value = value)
4. Clean Up (Optional)
If you don't need the group column anymore, you can drop it with dplyr::select:
a_final <- select(a_reshaped, -group)
Combine All Steps with Pipes
You can also chain all these operations together using the pipe operator (%>%) for cleaner code:
a_final <- a %>% gather(key = "variable", value = "value") %>% separate(variable, into = c("type", "group"), sep = 1) %>% spread(key = type, value = value) %>% select(-group)
Result
The final a_final DataFrame will have just two columns: x and y, with all the values from the original x1/x2/x3 and y1/y2/y3 columns stacked appropriately.
内容的提问来源于stack exchange,提问作者Minyi Han

