gather、reshape、cast等函数区别及长/宽数据转换实操咨询
Hey there! Let's break this down step by step—this data reshaping stuff can feel tricky at first, but once you wrap your head around the core concepts, it'll click.
Let's start with the basics to clear up those terms you mentioned:
- Long data: Your
dataframe is long data—each row represents a single observation, with a grouping variable (id) that repeats, and a value column (val) holding the measurements. Think of it as "tall and narrow". - Wide data: Your target
resframe is wide data—each group (A/B/C) becomes a column, and rows represent corresponding measurements across groups. Think "short and wide". - id variable: This is the variable that identifies a single "subject" or group. In your data,
id(A/B/C) is the id variable—it tells R which observations belong to the same group. - Time variable (or index variable): This is a variable that labels the position of each observation within its group. Your data doesn't have one yet, but we'll need to add one to make the conversion work—since each id has 10 values, we need a way to say "the 1st value of A matches the 1st value of B, etc."
All these tools handle data reshaping, but they belong to different R packages and have different syntax styles:
- Base R
reshape(): The old-school built-in function. It can do both long-to-wide and wide-to-long conversions, but its syntax is clunky with lots of parameters (likeidvar,timevar) that can be hard to remember. Good if you don't want to load extra packages, but not the most readable. reshape2package'scast()/melt(): A more user-friendly update to basereshape.melt()turns wide data into long data, anddcast()(for data frames) turns long data into wide data. It's more intuitive than basereshape, but it's been largely replaced by the tidyverse tools these days.- Tidyverse (
tidyr) functions:gather()(wide-to-long) andspread()(long-to-wide) were the original tidyverse reshaping tools, and now they've been superseded bypivot_longer()andpivot_wider()—these are the most readable and recommended for beginners, since their names clearly describe what they do, and they fit into the tidy data workflow.
Your data is missing that index variable we talked about—so first, we'll add a column to label which position each val is in for its group. Then we'll use different methods to reshape it:
Step 1: Add the index variable
First, let's add a row_num column to track the 1-10 position for each id:
data <- data.frame(id = c(rep("A", 10), rep("B", 10), rep("C", 10)), val = 1:30) data$row_num <- rep(1:10, 3) # Repeat 1-10 three times (once per id)
Method 1: Tidyverse's pivot_wider() (Recommended)
This is the most modern and readable approach:
library(tidyverse) res <- data %>% pivot_wider(names_from = id, # Use the 'id' values as column names values_from = val) # Fill the columns with 'val' values %>% select(-row_num) # Remove the index column if you don't need it
Method 2: reshape2's dcast()
If you're working with older code or prefer reshape2:
library(reshape2) res <- dcast(data, row_num ~ id, value.var = "val") res$row_num <- NULL # Remove the index column
The formula row_num ~ id means: "Use row_num to define rows, id to define columns, and fill with val values".
Method 3: Base R's reshape()
For when you don't want to load any packages:
res <- reshape(data, idvar = "row_num", # Variable to group rows by timevar = "id", # Variable to turn into columns direction = "wide") # Convert to wide format # Clean up column names (removes the "val." prefix) names(res) <- sub("val.", "", names(res)) res$row_num <- NULL # Remove the index column
- Base
reshape(): No package needed, but syntax is confusing and less readable. reshape2dcast()/melt(): More intuitive than base R, but less flexible than tidyverse tools now.- Tidyverse
pivot_wider()/pivot_longer(): Most user-friendly, readable, and integrates seamlessly with other tidyverse functions (likedplyrpipes). This is the go-to for most modern R users.
内容的提问来源于stack exchange,提问作者cgibbs_10

