针对特定分组修改dplyr中summarize()计算方法(附示例数据集)
First, let's complete your sample dataset with dummy count values so we can work through examples properly:
library(dplyr) df <- structure( list( x = structure( c(1L, 2L, 3L, 4L, 5L, 6L, 7L, 7L, 1L, 2L, 3L, 4L, 5L, 6L, 7L, 7L), .Label = c("1", "2", "3", "4", "5", "6", "Total"), class = "factor" ), y = structure( c(1L, 1L, 2L, 2L, 3L, 3L, 4L, 4L, 1L, 1L, 2L, 2L, 3L, 3L, 4L, 4L), .Label = c("7", "8", "9", "Total"), class = "factor" ), z = structure( c(1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L), .Label = c("10", "11"), class = "factor" ), count = c(56, 32, 45, 29, 61, 18, 241, 119, 56, 32, 45, 29, 61, 18, 241, 119) ), class = "data.frame", row.names = c(NA, -16L) )
Scenario 1: Summarize only non-"Total" groups
If you want to compute aggregated values (like sum of count) for groups where neither x nor y is "Total", filter those rows first before grouping and summarizing:
non_total_summary <- df %>% filter(x != "Total", y != "Total") %>% group_by(x, y, z) %>% summarize(total_count = sum(count), .groups = "drop") print(non_total_summary)
This gives you clean aggregated counts for every valid combination of x, y, and z excluding the precomputed total rows.
Scenario 2: Validate or recalculate "Total" groups
If you need to ensure the existing "Total" rows in your dataset are accurate, you can compute these totals programmatically and replace the existing values:
# Calculate true totals for each z value calculated_totals <- df %>% filter(x != "Total", y != "Total") %>% group_by(z) %>% summarize(calculated_total = sum(count), .groups = "drop") %>% mutate(x = "Total", y = "Total") # Combine with original non-total rows to create an updated dataset updated_df <- df %>% filter(x != "Total" | y != "Total") %>% bind_rows(calculated_totals) %>% arrange(x, y, z) print(updated_df)
This replaces the existing "Total" rows with values derived directly from your raw data, ensuring consistency.
Scenario 3: Custom summary logic for specific subgroups
If you want to apply different calculations to specific groups (e.g., sum for z="10" and average for z="11"), use conditional logic inside summarize():
custom_group_summary <- df %>% filter(x != "Total", y != "Total") %>% group_by(x, y) %>% summarize( z10_total = sum(count[z == "10"]), z11_avg = mean(count[z == "11"]), .groups = "drop" ) print(custom_group_summary)
This lets you tailor your summary exactly to the groups and metrics you care about.
If you had a specific grouping or calculation in mind that isn't covered here, feel free to share more details and I can adjust the code accordingly!
内容的提问来源于stack exchange,提问作者Clarinetist

