You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言sum函数中使用序列运算符":"按列名索引实现高效行求和?

Efficiently Summing Column Ranges in dplyr

Great question! Dealing with dynamic column ranges—especially those generated by pivot_wider—is a common pain point, so let's walk through cleaner, more robust solutions than your initial approach.

First, let's clarify why your rowwise() attempt failed: when using rowwise(), the : operator inside c() tries to create a sequence of values from the columns in that row, not select the columns themselves. Removing rowwise() works because dplyr can interpret the column range as a set of columns to sum across rows. But we can do better.

Solution 1: Use across() + rowSums() (Most Straightforward)

The across() function is built for targeting column ranges, and pairing it with rowSums() makes row-wise sums explicit and reliable:

df %>%
  mutate(Result_column = case_when(
    X == "A" ~ rowSums(across(X.1:X.5)),
    X == "B" ~ rowSums(across(X.4:X.5)),
    X == "C" ~ rowSums(across(X.3:X.56))
  )) %>%
  select(Result_column)

This avoids the ambiguity of sum(c(...)) and works consistently without needing rowwise().

Solution 2: Dynamic Column Ranges with num_range()

If you don't know the exact end of your column range (e.g., from pivot_wider creating variable columns), use num_range() to generate column names dynamically. This is perfect for consistent naming patterns like X.1, X.2, etc.:

# Get the maximum number from X.* column names
max_col_num <- max(as.integer(stringr::str_remove(names(df), "X\\.")))

df %>%
  mutate(Result_column = case_when(
    X == "A" ~ rowSums(across(num_range("X.", 1:5))),
    X == "B" ~ rowSums(across(num_range("X.", 4:5))),
    X == "C" ~ rowSums(across(num_range("X.", 3:max_col_num)))
  )) %>%
  select(Result_column)

Here, num_range("X.", 3:max_col_num) automatically selects all columns from X.3 to the last X.* column—no need to hardcode X.56!

Solution 3: rowwise() with c_across() (If You Need Row-Wise Context)

If you must use rowwise() (e.g., for other row-wise operations), replace c() with c_across(), which is designed to work with rowwise() to select columns per row:

df %>%
  rowwise() %>%
  mutate(Result_column = case_when(
    X == "A" ~ sum(c_across(X.1:X.5)),
    X == "B" ~ sum(c_across(X.4:X.5)),
    X == "C" ~ sum(c_across(X.3:X.56))
  )) %>%
  ungroup() %>% # Don't forget to ungroup afterward!
  select(Result_column)

c_across() correctly interprets the column range in a row-wise context, so sum() works as intended.

Key Takeaways

  • across() + rowSums() is the simplest approach for static column ranges.
  • num_range() is ideal for dynamic columns from pivot_wider or other variable datasets.
  • c_across() is your go-to if you need to stay in a rowwise() workflow.

All these methods are more maintainable than manually listing columns—especially as your dataset grows!

内容的提问来源于stack exchange,提问作者cn838

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 21:17:36