在R语言中按行匹配合并年份列表为单dataframe的问题
Hey there! Let's work through this problem together. The error you're seeing with do.call("cbind", examplelist) makes total sense—each dataframe in your list has a different number of rows (different State counts), so straight column binding can't align them properly. And since cbind.fill is deprecated, let's use modern, reliable approaches to get your desired wide-format dataframe.
Option 1: Use the Tidyverse (Recommended)
This method is clean and handles all the alignment and NA filling automatically. We'll first combine all your list elements into a single long-format dataframe, then reshape it to the wide format you want.
First, load the tidyverse package (if you haven't already):
library(tidyverse)
Then run these steps:
# Step 1: Combine all list elements into one long dataframe combined_long <- bind_rows(examplelist) # Step 2: Reshape to wide format, with State as rows and year-specific X/Y as columns Outcomewanted <- combined_long %>% pivot_wider( id_cols = State, # Keep State as the row identifier names_from = Year, # Use Year values to create new column names values_from = c(X, Y), # Spread X and Y values across the new columns names_glue = "{Year}_{.value}" # Format column names like "1971_X", "1971_Y" )
This works perfectly even if some list elements have multiple years (like your List[[2]] with 1972 and 1973). bind_rows will stack all rows together, and pivot_wider will automatically fill in NA for any State-Year combinations that don't exist in your data.
Option 2: Base R Alternative
If you prefer not to use the tidyverse, you can use a loop with merge to build your dataframe incrementally:
# Start with a dataframe containing States from the first list element (we'll expand this as we go) Outcomewanted <- data.frame(State = examplelist[[1]]$State) # Loop through each dataframe in your list for (df in examplelist) { # Get the unique year(s) in the current dataframe current_years <- unique(df$Year) # Process each year individually to build the correct columns for (year in current_years) { df_subset <- df[df$Year == year, c("State", "X", "Y")] colnames(df_subset)[2:3] <- paste0(year, "_", c("X", "Y")) # Merge with the existing result, keeping all States and filling gaps with NA Outcomewanted <- merge(Outcomewanted, df_subset, by = "State", all = TRUE) } }
This approach will also align all States correctly and fill NA where data is missing, just like the tidyverse method.
Why This Works
Both methods prioritize aligning rows by State instead of just stacking columns, which fixes the "differing number of rows" error you ran into earlier. They also handle the renaming of columns to include the year and automatically fill missing values with NA, exactly as you need.
内容的提问来源于stack exchange,提问作者JeffB

