You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用dplyr或reshape包展开R语言中聚合后的DataFrame,按水果计数生成重复行

Reshape Data to Repeat Rows per Fruit Instance in R

First, let's start with your sample dataframe so we can work with concrete, testable data:

library(tidyverse)

# Create the original dataframe
df <- tibble(
  Name = c("Tom", "John", "John", "John", "Alexa", "Alexa"),
  Year = c(2012, 2012, 2013, 2014, 2012, 2013),
  Apples = c(3, 0, 3, 5, 2, 2),
  Bananas = c(4, 1, 2, 3, 2, 1)
)

Using dplyr + tidyr (Tidyverse Approach)

This is the most straightforward and readable method for this task—it’s my go-to for data reshaping in R:

  1. Convert from wide to long format: Use pivot_longer() to turn the Apples and Bananas columns into a single Fruit column, with their counts stored in a new Count column.
  2. Filter out empty counts: We’ll drop rows where Count is 0 (like John’s 2012 Apples) since we don’t need rows for non-existent fruit instances.
  3. Repeat rows by count: Use uncount() to duplicate each row exactly Count times, which gives us one row per individual fruit.

Here’s the full code:

result <- df %>%
  pivot_longer(cols = c(Apples, Bananas), 
               names_to = "Fruit", 
               values_to = "Count") %>%
  filter(Count > 0) %>%  # Remove rows with 0 fruits
  uncount(Count) %>%     # Repeat rows to match the fruit count
  select(Name, Year, Fruit)  # Reorder columns to match your desired output

# Preview the first 10 rows to verify
head(result, 10)

Running this will produce exactly the output you outlined: every row represents one actual apple or banana that a person had in a given year.

Alternative: Using reshape Package

If you prefer to stick with the reshape package instead of the tidyverse, here’s an equivalent approach:

library(reshape)

# Melt the wide dataframe to long format
melted_df <- melt(df, id.vars = c("Name", "Year"), 
                  variable.name = "Fruit", 
                  value.name = "Count")

# Filter out 0-count rows and repeat rows by their count value
result_reshape <- melted_df[melted_df$Count > 0, ]
result_reshape <- result_reshape[rep(seq(nrow(result_reshape)), result_reshape$Count), ]

# Clean up the count column and reset row names
result_reshape$Count <- NULL
rownames(result_reshape) <- NULL

This achieves the same end result, though the tidyverse method is generally more concise and easier to follow for most modern R users.

内容的提问来源于stack exchange,提问作者aholtz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 06:39:07