如何在R语言sum函数中使用序列运算符":"按列名索引实现高效行求和?
Great question! Dealing with dynamic column ranges—especially those generated by pivot_wider—is a common pain point, so let's walk through cleaner, more robust solutions than your initial approach.
First, let's clarify why your rowwise() attempt failed: when using rowwise(), the : operator inside c() tries to create a sequence of values from the columns in that row, not select the columns themselves. Removing rowwise() works because dplyr can interpret the column range as a set of columns to sum across rows. But we can do better.
Solution 1: Use across() + rowSums() (Most Straightforward)
The across() function is built for targeting column ranges, and pairing it with rowSums() makes row-wise sums explicit and reliable:
df %>% mutate(Result_column = case_when( X == "A" ~ rowSums(across(X.1:X.5)), X == "B" ~ rowSums(across(X.4:X.5)), X == "C" ~ rowSums(across(X.3:X.56)) )) %>% select(Result_column)
This avoids the ambiguity of sum(c(...)) and works consistently without needing rowwise().
Solution 2: Dynamic Column Ranges with num_range()
If you don't know the exact end of your column range (e.g., from pivot_wider creating variable columns), use num_range() to generate column names dynamically. This is perfect for consistent naming patterns like X.1, X.2, etc.:
# Get the maximum number from X.* column names max_col_num <- max(as.integer(stringr::str_remove(names(df), "X\\."))) df %>% mutate(Result_column = case_when( X == "A" ~ rowSums(across(num_range("X.", 1:5))), X == "B" ~ rowSums(across(num_range("X.", 4:5))), X == "C" ~ rowSums(across(num_range("X.", 3:max_col_num))) )) %>% select(Result_column)
Here, num_range("X.", 3:max_col_num) automatically selects all columns from X.3 to the last X.* column—no need to hardcode X.56!
Solution 3: rowwise() with c_across() (If You Need Row-Wise Context)
If you must use rowwise() (e.g., for other row-wise operations), replace c() with c_across(), which is designed to work with rowwise() to select columns per row:
df %>% rowwise() %>% mutate(Result_column = case_when( X == "A" ~ sum(c_across(X.1:X.5)), X == "B" ~ sum(c_across(X.4:X.5)), X == "C" ~ sum(c_across(X.3:X.56)) )) %>% ungroup() %>% # Don't forget to ungroup afterward! select(Result_column)
c_across() correctly interprets the column range in a row-wise context, so sum() works as intended.
Key Takeaways
across()+rowSums()is the simplest approach for static column ranges.num_range()is ideal for dynamic columns frompivot_wideror other variable datasets.c_across()is your go-to if you need to stay in arowwise()workflow.
All these methods are more maintainable than manually listing columns—especially as your dataset grows!
内容的提问来源于stack exchange,提问作者cn838

