You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写符合DRY原则的简洁代码计算多列行均值?

DRY Solutions for Calculating Row Means of Replicate Columns

Great question! All that repetitive rowMeans code is exactly the kind of redundancy DRY (Don't Repeat Yourself) principles are designed to eliminate. Let's walk through clean, scalable solutions for both your specific 2-replicate case and the broader 8-replicate scenario you asked about.

First: Simplify Your 2-Replicate Code

Your current code repeats the same pattern for four different column groups—each time pairing xxx_rep1 and xxx_rep2 to create xxx_rep. Instead of writing each call manually, we can automate this using pattern matching and batch operations.

Base R (No External Packages)

This approach uses base R functions to detect replicate columns, extract their base names, and loop through them to calculate means:

# Extract base names by stripping "_rep1" or "_rep2" from column names
base_names <- unique(sub("_rep[12]$", "", names(dfg)[grepl("_rep[12]$", names(dfg))]))

# Loop through each base name and compute row means
for (name in base_names) {
  # Select the two replicate columns and calculate mean
  dfg[[paste0(name, "_rep")]] <- rowMeans(dfg[, paste0(name, "_rep", 1:2)], na.rm = TRUE)
}

This works even if you add more replicate groups later—no need to edit the loop itself.

Tidyverse (dplyr + stringr)

If you prefer a pipe-based workflow, the tidyverse makes this concise and readable:

library(dplyr)
library(stringr)

# Extract base names from columns ending with "_rep1"
base_names <- str_remove(names(dfg)[str_detect(names(dfg), "_rep1$")], "_rep1$")

# Use mutate() and across() to compute means for all groups at once
dfg <- dfg %>%
  mutate(
    across(
      all_of(base_names),
      ~ rowMeans(select(cur_data(), matches(paste0("^", .x, "_rep[12]$"))), na.rm = TRUE),
      .names = "{.col}_rep"
    )
  )

The .names argument automatically formats the new column names to match your original pattern (xxx_rep).

Scaling to 8 Columns (Replicates 1-8)

If you need to calculate row means for 8 replicate columns (e.g., xxx_rep1 through xxx_rep8), just adjust the pattern matching and selection logic:

Base R Adjustment

base_names <- unique(sub("_rep[1-8]$", "", names(dfg)[grepl("_rep[1-8]$", names(dfg))]))

for (name in base_names) {
  dfg[[paste0(name, "_rep")]] <- rowMeans(dfg[, paste0(name, "_rep", 1:8)], na.rm = TRUE)
}

Tidyverse Adjustment

base_names <- str_remove(names(dfg)[str_detect(names(dfg), "_rep1$")], "_rep1$")

dfg <- dfg %>%
  mutate(
    across(
      all_of(base_names),
      ~ rowMeans(select(cur_data(), matches(paste0("^", .x, "_rep[1-8]$"))), na.rm = TRUE),
      .names = "{.col}_rep"
    )
  )

Key Takeaway

The core idea is to identify the repeating pattern in your column names, then use batch operations (loops or across()) to apply the rowMeans calculation once instead of repeating it for every group. This keeps your code shorter, easier to maintain, and true to DRY principles.

内容的提问来源于stack exchange,提问作者user432797

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:40:22