如何使用expss包为计算得到的均值添加置信区间?
使用expss为均值添加置信区间
要在expss生成的表格中包含均值的置信区间,你可以通过tab_stat_fun_df()自定义统计函数来实现。以下是针对你的代码修改后的完整方案:
df <- data.frame(Group=c(1,1,1,1,1,1,1,0,0,0,0,0,0,0,0), Response=c(0,1,0,1,0,1,0,1,0,1,0,1,1,0,0), Value=c(10,20,17,14,15,16,17,18,19,10,11,12,13,14,15)) library(expss) # 自定义函数:计算均值及95%置信区间 mean_ci <- function(x) { # 过滤缺失值 x <- na.omit(x) n <- length(x) if(n == 0) { return(data.frame(Mean=NA, CI_Lower=NA, CI_Upper=NA)) } mean_val <- mean(x) se <- sd(x)/sqrt(n) # 计算95%置信区间的t分布临界值 ci_crit <- qt(0.975, df=n-1) ci_lower <- mean_val - se * ci_crit ci_upper <- mean_val + se * ci_crit # 返回包含统计量的数据框 data.frame(Mean=round(mean_val, 2), CI_Lower=round(ci_lower, 2), CI_Upper=round(ci_upper, 2)) } # 生成包含均值和置信区间的表格 df %>% tab_rows(Group) %>% tab_cols(Response) %>% tab_cells(Value) %>% tab_stat_fun_df(mean_ci) %>% # 调用自定义统计函数 tab_pivot()
代码说明:
mean_ci()函数:接收数值型向量,先过滤缺失值,再计算均值、标准误,基于t分布计算95%置信区间的上下限,最终返回包含这些统计量的数据框。tab_stat_fun_df():expss中用于自定义多统计量输出的接口,它支持返回数据框的函数,能一次性输出均值、置信区间上下限多个指标。- 若需调整置信水平,修改
qt()函数的概率参数即可(比如90%置信水平用0.95)。
如果希望将置信区间格式化为均值 (下限, 上限)的紧凑形式,可以使用以下修改后的函数:
mean_ci_formatted <- function(x) { x <- na.omit(x) n <- length(x) if(n == 0) { return(data.frame(Mean_CI=NA)) } mean_val <- mean(x) se <- sd(x)/sqrt(n) ci_crit <- qt(0.975, df=n-1) ci_lower <- mean_val - se * ci_crit ci_upper <- mean_val + se * ci_crit # 格式化为"均值 (下限, 上限)"的字符串 mean_ci_str <- sprintf("%.2f (%.2f, %.2f)", mean_val, ci_lower, ci_upper) data.frame(Mean_CI=mean_ci_str) } # 使用格式化后的函数生成表格 df %>% tab_rows(Group) %>% tab_cols(Response) %>% tab_cells(Value) %>% tab_stat_fun_df(mean_ci_formatted) %>% tab_pivot()
内容的提问来源于stack exchange,提问作者Mike
相关产品推荐
相关产品推荐

