You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr计算分组数据加权中位数的95%置信区间?

Solution: Weighted Median with 95% Confidence Intervals per Group

A robust, flexible way to compute weighted medians with 95% confidence intervals (CIs) for grouped data in dplyr is to use non-parametric bootstrapping. This approach avoids relying on undocumented functions or packages that lack weight support.

Step 1: Load Required Packages

We'll use dplyr for grouping, spatstat for weighted median calculation, and boot for bootstrapping:

library(dplyr)
library(spatstat)
library(boot)

Step 2: Define a Custom Function for Weighted Median + CI

This function computes the weighted median for a dataset and uses bootstrapping to generate 95% CIs:

get_weighted_median_ci <- function(df, val_col, wt_col, n_boot = 1000) {
  # Extract values and weights from the input data frame
  val <- df[[val_col]]
  wt <- df[[wt_col]]
  
  # Calculate the point estimate of the weighted median
  weighted_med <- weighted.median(val, wt)
  
  # Bootstrap function to resample observations with weight-proportional probability
  bootstrap_fn <- function(data, indices) {
    # Normalize weights to create sampling probabilities
    norm_weights <- data[[wt_col]] / sum(data[[wt_col]])
    # Sample rows with replacement using weight-based probabilities
    sampled_data <- data[sample(nrow(data), replace = TRUE, prob = norm_weights), ]
    # Compute weighted median for the bootstrap sample
    weighted.median(sampled_data[[val_col]], sampled_data[[wt_col]])
  }
  
  # Run the bootstrap procedure
  boot_results <- boot(data = df, statistic = bootstrap_fn, R = n_boot)
  
  # Extract the 95% percentile-based confidence interval
  ci <- boot.ci(boot_results, type = "perc")$percent[c(4, 5)]
  
  # Return results as a structured tibble
  tibble(
    weighted_median = weighted_med,
    ci_lower = ci[1],
    ci_upper = ci[2]
  )
}

Step 3: Apply to Grouped Data

Use dplyr::summarise() with cur_data() to apply the function to each group:

# Your original sample data
tst <- data.frame(group = rep(c(1:5), each = 100))
tst$val = runif(500) * tst$group
tst$wt = runif(500) * tst$val

# Compute weighted median and 95% CI per group
tst %>%
  group_by(group) %>%
  summarise(get_weighted_median_ci(cur_data(), "val", "wt"), .groups = "drop")

Key Details

  • Bootstrapping Logic: We resample observations from each group with probabilities proportional to their normalized weights, preserving the original weighted distribution. For each resample, we calculate the weighted median.
  • CI Calculation: The 95% CI is derived from the 2.5th and 97.5th percentiles of the bootstrap medians (percentile method), which is robust for non-normal distributions like the median.
  • dplyr Compatibility: cur_data() passes the current group's data frame to our custom function, ensuring seamless integration with grouped operations.

Notes

  • Adjust n_boot (default: 1000) for more stable CIs (higher values increase computation time but improve reliability).
  • To use a different CI type (e.g., basic or BCa), modify the type argument in boot.ci().

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 03:45:38