如何使用dplyr或reshape包展开R语言中聚合后的DataFrame,按水果计数生成重复行
Reshape Data to Repeat Rows per Fruit Instance in R
First, let's start with your sample dataframe so we can work with concrete, testable data:
library(tidyverse) # Create the original dataframe df <- tibble( Name = c("Tom", "John", "John", "John", "Alexa", "Alexa"), Year = c(2012, 2012, 2013, 2014, 2012, 2013), Apples = c(3, 0, 3, 5, 2, 2), Bananas = c(4, 1, 2, 3, 2, 1) )
Using dplyr + tidyr (Tidyverse Approach)
This is the most straightforward and readable method for this task—it’s my go-to for data reshaping in R:
- Convert from wide to long format: Use
pivot_longer()to turn theApplesandBananascolumns into a singleFruitcolumn, with their counts stored in a newCountcolumn. - Filter out empty counts: We’ll drop rows where
Countis 0 (like John’s 2012 Apples) since we don’t need rows for non-existent fruit instances. - Repeat rows by count: Use
uncount()to duplicate each row exactlyCounttimes, which gives us one row per individual fruit.
Here’s the full code:
result <- df %>% pivot_longer(cols = c(Apples, Bananas), names_to = "Fruit", values_to = "Count") %>% filter(Count > 0) %>% # Remove rows with 0 fruits uncount(Count) %>% # Repeat rows to match the fruit count select(Name, Year, Fruit) # Reorder columns to match your desired output # Preview the first 10 rows to verify head(result, 10)
Running this will produce exactly the output you outlined: every row represents one actual apple or banana that a person had in a given year.
Alternative: Using reshape Package
If you prefer to stick with the reshape package instead of the tidyverse, here’s an equivalent approach:
library(reshape) # Melt the wide dataframe to long format melted_df <- melt(df, id.vars = c("Name", "Year"), variable.name = "Fruit", value.name = "Count") # Filter out 0-count rows and repeat rows by their count value result_reshape <- melted_df[melted_df$Count > 0, ] result_reshape <- result_reshape[rep(seq(nrow(result_reshape)), result_reshape$Count), ] # Clean up the count column and reset row names result_reshape$Count <- NULL rownames(result_reshape) <- NULL
This achieves the same end result, though the tidyverse method is generally more concise and easier to follow for most modern R users.
内容的提问来源于stack exchange,提问作者aholtz
相关产品推荐
相关产品推荐

