在R语言中计算指定Count值的累计百分比技术问题
Hey there! Let's fix this up for you. The issue with your original code is that it's calculating the cumulative sum of the Count values themselves (not the number of rows per Count group) divided by the total sum of all Count values. That's why you're getting those small incremental percentages instead of the group-level cumulative totals you want.
Here's how to get the cumulative percentages you're expecting (50% for Count=1, 80% for Count=2, 100% for Count=3):
Method 1: Summarize first, then join back
First, we'll calculate the number of rows per Count group, compute the cumulative percentage for each group, then merge that data back into your original tibble so every row gets the correct cumulative percentage for its Count value:
library(dplyr) # Step 1: Create a summary table with cumulative percentages per Count count_cum <- d2 %>% count(Count) %>% # Get number of rows for each Count mutate(cum_percent = cumsum(n) / sum(n) * 100) # Calculate cumulative percentage # Step 2: Join this summary back to the original data d2_result <- d2 %>% left_join(count_cum, by = "Count")
Method 2: Window functions in one pipeline
If you prefer a single pipeline without a separate summary table, you can use window functions to compute group sizes and cumulative totals directly:
d2_result <- d2 %>% arrange(Count) %>% # Ensure we're processing Count values in order group_by(Count) %>% mutate(group_size = n()) %>% # Number of rows in each Count group ungroup() %>% mutate( cum_total = cumsum(group_size), # Cumulative number of rows up to each Count cum_percent = cum_total / sum(group_size) * 100 ) %>% select(-group_size, -cum_total) # Clean up extra columns if needed
What this does
- Both methods first calculate how many rows belong to each
Countvalue (5 rows for Count=1, 3 for Count=2, 2 for Count=3). - Then they compute the cumulative sum of these row counts (5, 5+3=8, 8+2=10) and divide by the total number of rows (10) to get the cumulative percentages (50%, 80%, 100%).
- Finally, they attach these percentages back to every row in the original data, so all rows with Count=1 get 50%, Count=2 get 80%, etc.
Example output
Your resulting tibble will look like this:
# A tibble: 10 x 3 ID Count cum_percent <int> <dbl> <dbl> 1 1 1 50 2 1 1 50 3 1 1 50 4 1 1 50 5 1 1 50 6 2 2 80 7 2 2 80 8 2 2 80 9 3 3 100 10 3 3 100
内容的提问来源于stack exchange,提问作者Omry Atia

