使用dplyr包通过列索引按子组计算加权均值
Sure thing! Let's walk through how to pull this off with dplyr using column indices instead of column names. It’s straightforward once you leverage dplyr’s built-in tools for referencing columns by position.
First, let’s fix up the example data loading code (the original had formatting issues that would cause errors):
data <- read.table(text = 'obs income education type weight 1 1000 A blue 10 2 2000 B yellow 1 3 1500 B blue 5 4 2000 A yellow 2 5 3000 B yellow 2', header = TRUE)
Step 1: Map Columns to Indices
First, let’s clarify which columns correspond to which indices in your data:
- Target variable (we want the weighted mean of this):
income→ Column 2 - Grouping variables:
education(Column 3) andtype(Column 4) - Weight variable:
weight→ Column 5
Step 2: Group by Column Indices & Calculate Weighted Mean
We’ll use across() inside group_by() to specify grouping columns by index, and dplyr’s .data pronoun to reference the target and weight columns by position.
Option 1: Add Weighted Mean as a New Column (Retain All Rows)
If you want to keep every row and add the group-level weighted mean to each entry:
library(dplyr) result_with_all_rows <- data %>% group_by(across(c(3, 4))) %>% # Group by columns 3 and 4 mutate( weighted_avg_income = weighted.mean( x = .data[[2]], # Target variable: column 2 (income) w = .data[[5]], # Weight variable: column 5 (weight) na.rm = TRUE # Handle missing values safely ) ) %>% ungroup()
Option 2: Summarize to Get One Row Per Group
If you only want the aggregated weighted mean for each group:
grouped_summary <- data %>% group_by(across(c(3, 4))) %>% summarize( weighted_avg_income = weighted.mean(.data[[2]], .data[[5]], na.rm = TRUE), .groups = "drop" # Clean up grouping structure after summarizing )
Running this will give you a tidy tibble with one row per education/type group and its corresponding weighted average income:
# A tibble: 4 × 3 education type weighted_avg_income <chr> <chr> <dbl> 1 A blue 1000 2 A yellow 2000 3 B blue 1500 4 B yellow 2667.
Bonus: Dynamic Column Indices
If you want to store your column indices in variables (great for reusable code), use all_of() inside across() to reference them:
# Define indices as variables target_col <- 2 group_cols <- c(3, 4) weight_col <- 5 dynamic_result <- data %>% group_by(across(all_of(group_cols))) %>% summarize( weighted_avg = weighted.mean(.data[[target_col]], .data[[weight_col]], na.rm = TRUE), .groups = "drop" )
This works because all_of() tells dplyr to treat the group_cols vector as literal column positions, rather than trying to match variable names in the environment.
内容的提问来源于stack exchange,提问作者Elixterra

