R语言中如何将tapply生成的Class数组转换为数据帧?
Hey Luis, let's work through this problem step by step! I see a couple of small issues in your current code, plus we can make this way more efficient instead of manually binding each year's data.
First, let's fix your existing code
The main problems are:
- A typo in your
table()call (missing a closing parenthesis) - Incorrect indexing for the 2015 entry (you used
a["2015", "1st Qu."]instead of the list-stylea[["2015"]][["1st Qu."]]) - A variable name mismatch when setting column names (you referenced
losstrokebut your data frame isxdata)
Here's the corrected version of your original code:
# Fix the table call first (add closing parenthesis) table(Data$year, useNA=c("always")) # Your tapply code is fine a <- tapply(Data$x, Data$year, summary) # Correct the indexing and variable name xdata <- as.data.frame(rbind( x2014 <- c(2014, a[["2014"]][["1st Qu."]], a[["2014"]][["Median"]], a[["2014"]][["3rd Qu."]]), x2015 <- c(2015, a[["2015"]][["1st Qu."]], a[["2015"]][["Median"]], a[["2015"]][["3rd Qu."]]), x2016 <- c(2016, a[["2016"]][["1st Qu."]], a[["2016"]][["Median"]], a[["2016"]][["3rd Qu."]]) )) colnames(xdata) <- c("Year", "Q1", "Median", "Q3") rm(a, x2014, x2015, x2016)
A better, automated approach (no manual binding!)
Manually writing out each year isn't scalable if you add more years later. Instead, we can convert that a list directly into a data frame:
Using base R:
# Convert the list of summaries into a data frame summary_df <- as.data.frame(do.call(rbind, a)) # Add the Year column using the list's names summary_df$Year <- rownames(summary_df) # Extract only the columns we need and rename them final_df <- summary_df[, c("Year", "1st Qu.", "Median", "3rd Qu.")] colnames(final_df) <- c("Year", "Q1", "Median", "Q3") # Optional: Convert Year to numeric if needed final_df$Year <- as.numeric(final_df$Year)
Using dplyr (tidyverse style, even cleaner):
If you're open to using the dplyr package, we can skip the tapply step entirely and summarize directly from your original Data frame—this is more readable and less error-prone:
library(dplyr) final_df <- Data %>% group_by(year) %>% summarize( Q1 = quantile(x, 0.25, na.rm = TRUE), # Add na.rm=TRUE if you have missing values Median = median(x, na.rm = TRUE), Q3 = quantile(x, 0.75, na.rm = TRUE) )
Why your original code broke
When you use tapply() with summary, it returns a named list (not a matrix or data frame). Each element of the list is a named numeric vector containing the summary stats for one year. That means you can't use matrix-style indexing like a["2015", "1st Qu."]—you have to use list indexing: a[["2015"]] to get the 2015 summary vector, then [["1st Qu."]] to pull the specific stat from that vector.
内容的提问来源于stack exchange,提问作者Luis Labán

