在R语言中计算分组求和、总计及排名并生成新列
Solution for Grouped Row Sums, Total, and Ranking in R
Got it, let's break this down so it works even with your 200-column dataset—no need to manually list columns! Here's a clean, scalable approach using both tidyverse (dplyr) and base R, whichever you prefer:
Using dplyr (Tidyverse)
This method is super readable, especially for large datasets:
# Load the package (install first if needed: install.packages("dplyr")) library(dplyr) # Your sample data df <- read.table(text="Q1a Q2a Q3b Q4c Q5a Q6c Q7b 1 2 4 2 2 0 1 3 2 1 2 2 1 1 4 3 2 1 1 1 1", h=T) # Add grouped sums, total, and rank result_df <- df %>% # Calculate row sums for each group (a, b, c) by column suffix mutate( a = rowSums(select(., ends_with("a"))), b = rowSums(select(., ends_with("b"))), c = rowSums(select(., ends_with("c"))), # Total is sum of a, b, c Total = a + b + c, # Rank by Total (descending, same totals get same rank) Rank = min_rank(desc(Total)) ) # View the result (matches your expected output) result_df[, c(colnames(df), "a", "b", "c", "Total", "Rank")]
Using Base R (No Packages Needed)
If you prefer not to load external packages, this works just as well:
# Your sample data df <- read.table(text="Q1a Q2a Q3b Q4c Q5a Q6c Q7b 1 2 4 2 2 0 1 3 2 1 2 2 1 1 4 3 2 1 1 1 1", h=T) # Calculate grouped row sums using regex to match column suffixes df$a <- rowSums(df[, grepl("a$", colnames(df))]) df$b <- rowSums(df[, grepl("b$", colnames(df))]) df$c <- rowSums(df[, grepl("c$", colnames(df))]) # Add Total and Rank df$Total <- df$a + df$b + df$c # Use rank() with ties.method="min" to get same rank for equal totals df$Rank <- rank(-df$Total, ties.method = "min") # View the final table df
Key Notes:
- Both methods scale perfectly to 200 columns—no need to adjust code for more columns, since we're using pattern matching (
ends_withor regex) instead of hardcoding column names. - The ranking uses min rank (same totals get the same rank, which matches your sample output). If you want different ranking behavior (like dense rank), just swap
min_rank()fordense_rank()in dplyr, or adjust theties.methodin base R.
内容的提问来源于stack exchange,提问作者user9285150
相关产品推荐
相关产品推荐

