如何基于按年龄统计的App用户数计算年龄中位数与IQR?
计算加权年龄中位数与IQR的解决方案
你的数据是分组汇总后的形式(年龄为分组项,用户数为对应权重),普通的中位数/IQR函数无法直接适用,需要通过加权分位数的逻辑计算。以下是两种实用的R实现方案:
方案一:使用Hmisc包(推荐,成熟稳定)
Hmisc包提供了专门的加权分位数计算函数,能直接输出Q1、中位数、Q3和IQR。
步骤1:安装并加载依赖包
# 首次运行需安装包 install.packages("Hmisc") # 加载包 library(Hmisc)
步骤2:加载你的数据
df <- structure(list( `Row Labels` = c(0, 1, 2, 3, 4, 5), b_21 = c(35, 5, 5, 4, 3, 4), b_22 = c(79, 68, 87, 66, 44, 36), q_21 = c(56, 34, 27, 10, 12, 10), q_22 = c(71, 64, 45, 25, 27, 16), wc_21 = c(1226, 159, 98, 56, 48, 39), wc_22 = c(1127, 120, 73, 66, 38, 32), total = c(2594, 450, 335, 227, 172, 137)), row.names = c(NA, -6L), class = c("tbl_df", "tbl", "data.frame") )
步骤3:定义统计计算函数并批量运行
# 定义函数:输入某年份的用户数列,输出加权分位数统计 calc_weighted_stats <- function(weights) { ages <- df$`Row Labels` # 计算25%、50%、75%分位数 quantiles <- wtd.quantile(ages, weights = weights, probs = c(0.25, 0.5, 0.75)) # 整理结果 data.frame( Q1 = quantiles[1], Median = quantiles[2], Q3 = quantiles[3], IQR = quantiles[3] - quantiles[1] ) } # 选择所有年份列(排除第一列年龄) year_cols <- colnames(df)[-1] # 批量计算各年份统计量并转为易读格式 result_df <- do.call(rbind, lapply(df[year_cols], calc_weighted_stats)) rownames(result_df) <- year_cols # 查看结果 print(result_df)
方案二:自定义函数(无依赖,适合个性化需求)
如果不想引入第三方包,可以手动实现加权分位数逻辑:
# 自定义加权中位数函数 weighted_median <- function(x, w) { ord <- order(x) x_sorted <- x[ord] w_sorted <- w[ord] cum_w_ratio <- cumsum(w_sorted) / sum(w_sorted) # 找到第一个累计权重占比≥50%的年龄 x_sorted[which(cum_w_ratio >= 0.5)[1]] } # 自定义加权IQR函数(Q3-Q1) weighted_iqr <- function(x, w) { ord <- order(x) x_sorted <- x[ord] w_sorted <- w[ord] cum_w_ratio <- cumsum(w_sorted) / sum(w_sorted) q1_pos <- which(cum_w_ratio >= 0.25)[1] q3_pos <- which(cum_w_ratio >= 0.75)[1] x_sorted[q3_pos] - x_sorted[q1_pos] } # 批量计算各年份的中位数和IQR custom_result <- data.frame( Median = sapply(df[year_cols], function(col) weighted_median(df$`Row Labels`, col)), IQR = sapply(df[year_cols], function(col) weighted_iqr(df$`Row Labels`, col)) ) rownames(custom_result) <- year_cols print(custom_result)
内容的提问来源于stack exchange,提问作者edward schenck
相关产品推荐
相关产品推荐

