如何按变量类别计算行均值?R语言数据框处理问题
Hey there! Let's fix this row mean problem for your dataframe df—I’ve run into similar hiccups before, so I know exactly where you might be tripping up. Here are a few solid, error-free ways to get those category means added as new columns:
rowMeans()(简单直接) If your variable names are straightforward (like weight1, weight2, weight3), this is the quickest fix. Just make sure to handle missing values if you have any:
# 计算体重行均值,若不需要忽略NA则去掉na.rm=TRUE df$mean_weight <- rowMeans(df[, c("weight1", "weight2", "weight3")], na.rm = TRUE) # 身高类变量同理 df$mean_height <- rowMeans(df[, c("height1", "height2", "height3")], na.rm = TRUE)
A common reason for your original error? Likely a typo in column names, or accidentally selecting non-numeric columns. Double-check that the columns you’re targeting are all numeric with str(df) if you’re unsure.
If you prefer tidyverse syntax, across() is perfect for grouping variables by name patterns—great if you have more than 3 variables per category:
First load the package:
library(dplyr)
Then pick one of these two approaches:
- Pattern-matching with
across(): Automatically grabs all columns starting with your category prefix
df <- df %>% mutate( mean_weight = rowMeans(across(starts_with("weight")), na.rm = TRUE), mean_height = rowMeans(across(starts_with("height")), na.rm = TRUE) )
You can also use regex for stricter matching, like matches("^weight\\d+$") to only catch columns with "weight" followed by numbers.
- Row-wise calculation: More intuitive if you want to explicitly list variables
df <- df %>% rowwise() %>% mutate( mean_weight = mean(c(weight1, weight2, weight3), na.rm = TRUE), mean_height = mean(c(height1, height2, height3), na.rm = TRUE) ) %>% ungroup() # 别忘了取消行分组,避免后续操作变慢
If you have tons of variable groups (not just weight and height), a small loop will save you repetitive code:
# 定义所有类别前缀 categories <- c("weight", "height") # 可以添加更多,比如"age", "bp"等 for (cat in categories) { # 匹配该类别下的所有列 target_cols <- grep(paste0("^", cat), names(df), value = TRUE) # 生成新列名 mean_col_name <- paste0("mean_", cat) # 计算并添加行均值列 df[[mean_col_name]] <- rowMeans(df[, target_cols], na.rm = TRUE) }
Quick troubleshooting note: If your original error mentioned "replacement has X rows, data has Y rows", it means the mean vector you tried to add doesn’t match the number of rows in df. Double-check that you’re selecting the right columns (no typos!) and that they’re all numeric.
内容的提问来源于stack exchange,提问作者Bertrand

