使用infer包检验标准差时,模拟零分布未围绕零假设中心的疑问
问题原因及解决方案
你的核心误区在于infer包的bootstrap生成逻辑不会自动根据原假设的标准差调整数据:
当你使用generate(type = "bootstrap")时,它只是对原始数据(sd=1)做有放回重抽样,完全没有用到hypothesise()里指定的sigma=5参数——这个参数目前在infer的bootstrap工作流中并没有被用来调整样本数据,所以生成的bootstrap样本依然基于原始数据的标准差,计算出的统计量自然围绕原始sd=1,而非原假设的5。
而你提供的参数检验代码,是基于标准差检验的理论分布:对于原假设$H_0: \sigma=5$,统计量$\frac{(n-1)s2}{\sigma_02}$服从自由度为$n-1$的卡方分布,你通过逆变换生成的样本分布是围绕5的,这和infer的bootstrap逻辑完全不同。
修正方法:手动按原假设缩放数据后再bootstrap
要让bootstrap零分布围绕原假设的sigma=5,需要先将原始数据调整为符合原假设的分布,再进行bootstrap抽样:
library(tidyverse) library(infer) data <- tibble(x = rnorm(100, mean = 2, sd = 1)) # 先按原假设缩放数据:标准化后乘以原假设sigma,再加回原始均值 adjusted_data <- data %>% mutate(x_adjusted = (x - mean(x))/sd(x) * 5 + mean(x)) # 用调整后的数据生成bootstrap零分布 adjusted_data %>% specify(response = x_adjusted) %>% hypothesise(null = "point", sigma = 5) %>% generate(reps = 1000, type = "bootstrap") %>% calculate("sd") %>% visualise()
这样生成的bootstrap分布就会围绕5,和你参数检验的结果接近。
验证对比
你可以把修正后的bootstrap分布和参数检验分布放在一起对比:
# 生成修正后的bootstrap统计量 boot_stats <- adjusted_data %>% specify(response = x_adjusted) %>% hypothesise(null = "point", sigma = 5) %>% generate(reps = 1000, type = "bootstrap") %>% calculate("sd") %>% pull(stat) # 生成参数检验统计量 param_stats <- sqrt(rchisq(1000, 99)/(99)*(5)^2) # 对比绘图 tibble( type = rep(c("bootstrap", "parametric"), each = 1000), stat = c(boot_stats, param_stats) ) %>% ggplot(aes(x = stat, fill = type)) + geom_histogram(bins = 15, alpha = 0.7, position = "identity", color = "white") + labs(x = "stat", y = "count") + theme_minimal()
可以看到两者分布基本重合,符合预期。
内容的提问来源于stack exchange,提问作者2p4olE3xjEQJ
相关产品推荐
相关产品推荐

