You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr包通过列索引按子组计算加权均值

Calculate Weighted Mean by Groups Using Column Indices in dplyr

Sure thing! Let's walk through how to pull this off with dplyr using column indices instead of column names. It’s straightforward once you leverage dplyr’s built-in tools for referencing columns by position.

First, let’s fix up the example data loading code (the original had formatting issues that would cause errors):

data <- read.table(text = 'obs income education type weight
1 1000 A blue 10
2 2000 B yellow 1
3 1500 B blue 5
4 2000 A yellow 2
5 3000 B yellow 2', header = TRUE)

Step 1: Map Columns to Indices

First, let’s clarify which columns correspond to which indices in your data:

  • Target variable (we want the weighted mean of this): income → Column 2
  • Grouping variables: education (Column 3) and type (Column 4)
  • Weight variable: weight → Column 5

Step 2: Group by Column Indices & Calculate Weighted Mean

We’ll use across() inside group_by() to specify grouping columns by index, and dplyr’s .data pronoun to reference the target and weight columns by position.

Option 1: Add Weighted Mean as a New Column (Retain All Rows)

If you want to keep every row and add the group-level weighted mean to each entry:

library(dplyr)

result_with_all_rows <- data %>%
  group_by(across(c(3, 4))) %>%  # Group by columns 3 and 4
  mutate(
    weighted_avg_income = weighted.mean(
      x = .data[[2]],  # Target variable: column 2 (income)
      w = .data[[5]],  # Weight variable: column 5 (weight)
      na.rm = TRUE     # Handle missing values safely
    )
  ) %>%
  ungroup()

Option 2: Summarize to Get One Row Per Group

If you only want the aggregated weighted mean for each group:

grouped_summary <- data %>%
  group_by(across(c(3, 4))) %>%
  summarize(
    weighted_avg_income = weighted.mean(.data[[2]], .data[[5]], na.rm = TRUE),
    .groups = "drop"  # Clean up grouping structure after summarizing
  )

Running this will give you a tidy tibble with one row per education/type group and its corresponding weighted average income:

# A tibble: 4 × 3
  education type  weighted_avg_income
  <chr>     <chr>               <dbl>
1 A         blue                1000
2 A         yellow              2000
3 B         blue                1500
4 B         yellow              2667.

Bonus: Dynamic Column Indices

If you want to store your column indices in variables (great for reusable code), use all_of() inside across() to reference them:

# Define indices as variables
target_col <- 2
group_cols <- c(3, 4)
weight_col <- 5

dynamic_result <- data %>%
  group_by(across(all_of(group_cols))) %>%
  summarize(
    weighted_avg = weighted.mean(.data[[target_col]], .data[[weight_col]], na.rm = TRUE),
    .groups = "drop"
  )

This works because all_of() tells dplyr to treat the group_cols vector as literal column positions, rather than trying to match variable names in the environment.

内容的提问来源于stack exchange,提问作者Elixterra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:35:13