You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言如何在相对频率直方图上叠加尺度匹配的正态分布曲线

问题根源

你当前代码手动将直方图的密度值替换为 组内计数/总计数 *100,纵轴实际是百分比形式的组相对频率,而直接调用dnorm()生成的正态曲线值是概率密度尺度(全区间积分和为1),二者尺度不匹配,所以曲线无法对齐。

换算规则

要让正态曲线适配当前纵轴,需要对dnorm()的输出做两步换算:

  1. 乘以直方图的组距:概率密度 * 组距 = 对应组的理论概率
  2. 乘以100:将概率转换为和纵轴一致的百分比相对频率

换算公式:
适配纵轴的曲线y值 = dnorm(横坐标x, 正态均值, 正态标准差) * 直方图组距 * 100

根据中心极限定理,不同样本量下样本均值对应的正态分布参数不同:

  • 均值统一为均匀分布的期望10
  • 标准差为 sqrt(均匀分布方差/对应样本量n)
修正后完整代码
set.seed(1099)

N <- 1520
n_1 <- 4
n_2 <- 30
n_3 <- 76
Valor_esperado = (8 + 12)/2
Variancia = (12-8)^2/12

Amostra_1 <- matrix( runif(N*n_1,min = 8,max = 12), nrow = n_1)
Amostra_2 <- matrix( runif(N*n_2,min = 8,max = 12), nrow = n_2)
Amostra_3 <- matrix( runif(N*n_3,min = 8,max = 12), nrow = n_3)


media_1 <- colMeans(Amostra_1)
media_2 <- colMeans(Amostra_2)
media_3 <- colMeans(Amostra_3)


Amostra_1 <- as.numeric(unlist(media_1))
Amostra_2 <- as.numeric(unlist(media_2))
Amostra_3 <- as.numeric(unlist(media_3))

par(mfrow=c(1,3)) # 三个图并排展示方便对比

# 绘制n=4的直方图+匹配尺度的正态曲线
h <- hist(Amostra_1, plot=FALSE)
bin_width <- diff(h$breaks)[1] # 自动提取直方图组距
h$density = h$counts/sum(h$counts) * 100
plot(h, main="n = 4",
     xlab = NULL,
     ylab="Frequência Relativa",
     col="blue",
     freq=FALSE)
curve_x <- seq(min(h$breaks), max(h$breaks), length.out = 100)
curve_y <- dnorm(curve_x, 
                 mean = Valor_esperado, 
                 sd = sqrt(Variancia/n_1)) * bin_width * 100
lines(curve_x, curve_y, col = "black", lwd = 2)

# 绘制n=30的直方图+匹配尺度的正态曲线
h <- hist(Amostra_2, plot=FALSE)
bin_width <- diff(h$breaks)[1]
h$density = h$counts/sum(h$counts) * 100
plot(h, main="n = 30",
     xlab = NULL,
     ylab="Frequência Relativa",
     col="red",
     freq=FALSE)
curve_x <- seq(min(h$breaks), max(h$breaks), length.out = 100)
curve_y <- dnorm(curve_x, 
                 mean = Valor_esperado, 
                 sd = sqrt(Variancia/n_2)) * bin_width * 100
lines(curve_x, curve_y, col = "black", lwd = 2)

# 绘制n=76的直方图+匹配尺度的正态曲线
h <- hist(Amostra_3, plot=FALSE)
bin_width <- diff(h$breaks)[1]
h$density = h$counts/sum(h$counts) * 100
plot(h, main="n = 76",
     xlab = NULL,
     ylab="Frequência Relativa",
     col="yellow",
     freq=FALSE)
curve_x <- seq(min(h$breaks), max(h$breaks), length.out = 100)
curve_y <- dnorm(curve_x, 
                 mean = Valor_esperado, 
                 sd = sqrt(Variancia/n_3)) * bin_width * 100
lines(curve_x, curve_y, col = "black", lwd = 2)

内容的提问来源于stack exchange,提问作者Nome_aleatorio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 20:45:36