You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于按年龄统计的App用户数计算年龄中位数与IQR?

计算加权年龄中位数与IQR的解决方案

你的数据是分组汇总后的形式(年龄为分组项,用户数为对应权重),普通的中位数/IQR函数无法直接适用,需要通过加权分位数的逻辑计算。以下是两种实用的R实现方案:


方案一:使用Hmisc包(推荐,成熟稳定)

Hmisc包提供了专门的加权分位数计算函数,能直接输出Q1、中位数、Q3和IQR。

步骤1:安装并加载依赖包

# 首次运行需安装包
install.packages("Hmisc")
# 加载包
library(Hmisc)

步骤2:加载你的数据

df <- structure(list(
  `Row Labels` = c(0, 1, 2, 3, 4, 5), 
  b_21 = c(35, 5, 5, 4, 3, 4), 
  b_22 = c(79, 68, 87, 66, 44, 36),
  q_21 = c(56, 34, 27, 10, 12, 10), 
  q_22 = c(71, 64, 45, 25, 27, 16), 
  wc_21 = c(1226, 159, 98, 56, 48, 39), 
  wc_22 = c(1127, 120, 73, 66, 38, 32), 
  total = c(2594, 450, 335, 227, 172, 137)), 
  row.names = c(NA, -6L), 
  class = c("tbl_df", "tbl", "data.frame")
)

步骤3:定义统计计算函数并批量运行

# 定义函数:输入某年份的用户数列,输出加权分位数统计
calc_weighted_stats <- function(weights) {
  ages <- df$`Row Labels`
  # 计算25%、50%、75%分位数
  quantiles <- wtd.quantile(ages, weights = weights, probs = c(0.25, 0.5, 0.75))
  # 整理结果
  data.frame(
    Q1 = quantiles[1],
    Median = quantiles[2],
    Q3 = quantiles[3],
    IQR = quantiles[3] - quantiles[1]
  )
}

# 选择所有年份列(排除第一列年龄)
year_cols <- colnames(df)[-1]
# 批量计算各年份统计量并转为易读格式
result_df <- do.call(rbind, lapply(df[year_cols], calc_weighted_stats))
rownames(result_df) <- year_cols

# 查看结果
print(result_df)

方案二:自定义函数(无依赖,适合个性化需求)

如果不想引入第三方包,可以手动实现加权分位数逻辑:

# 自定义加权中位数函数
weighted_median <- function(x, w) {
  ord <- order(x)
  x_sorted <- x[ord]
  w_sorted <- w[ord]
  cum_w_ratio <- cumsum(w_sorted) / sum(w_sorted)
  # 找到第一个累计权重占比≥50%的年龄
  x_sorted[which(cum_w_ratio >= 0.5)[1]]
}

# 自定义加权IQR函数(Q3-Q1)
weighted_iqr <- function(x, w) {
  ord <- order(x)
  x_sorted <- x[ord]
  w_sorted <- w[ord]
  cum_w_ratio <- cumsum(w_sorted) / sum(w_sorted)
  q1_pos <- which(cum_w_ratio >= 0.25)[1]
  q3_pos <- which(cum_w_ratio >= 0.75)[1]
  x_sorted[q3_pos] - x_sorted[q1_pos]
}

# 批量计算各年份的中位数和IQR
custom_result <- data.frame(
  Median = sapply(df[year_cols], function(col) weighted_median(df$`Row Labels`, col)),
  IQR = sapply(df[year_cols], function(col) weighted_iqr(df$`Row Labels`, col))
)
rownames(custom_result) <- year_cols

print(custom_result)

内容的提问来源于stack exchange,提问作者edward schenck

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 17:37:26