求助:如何平滑直方图上的正态曲线?geom_smooth是否适用?
如何平滑ggplot中叠加的正态曲线?

我已经手动选择数据范围,用ggplot绘制了叠加两条正态曲线的直方图,但刚接触RStudio,不知道怎么平滑曲线。想问下geom_smooth()是否适用,或者有没有其他解决方法?
当前代码
#calculate normal curve and set interval data_m1 <- data %>% filter(masses_kDa >= 0 & masses_kDa <= 150) #change range here data_m2 <- data %>% filter(masses_kDa >= 300 & masses_kDa <= 900) #change range here #calculate mean and standard deviation mean_m1 <- mean(data_m1$masses_kDa) sd_m1 <- sd(data_m1$masses_kDa) mean_m2 <- mean(data_m2$masses_kDa) sd_m2 <- sd(data_m2$masses_kDa) library(ggplot2) # Calculate the bin width and the total number of data points bin_width <- 50 total_points_m1 <- length(data_m1$masses_kDa) total_points_m2 <- length(data_m2$masses_kDa) # Plot the histogram ggplot(cleaned_data, aes(x=masses_kDa)) + geom_histogram(color='transparent', fill='red', alpha = 0.5, bins= 50) + # Adjust the number of bins here labs(x=' ', y=' ', title=' ') + # Add the normal curves to the histogram stat_function(fun = function(x) dnorm(x, mean = mean_m1, sd = sd_m1) * bin_width * total_points_m1, col= '#8B0000') + stat_function(fun = function(x) dnorm(x, mean = mean_m2, sd = sd_m2) * bin_width * total_points_m2, col= '#8B0000') + # Add the mean lines to the histogram geom_vline(aes(xintercept = mean_m1), linetype="dashed", color = 'black') + geom_vline(aes(xintercept = mean_m2), linetype="dashed", color = 'black') + # Ensure there are no gaps between the graph and axis scale_x_continuous(expand = c(0,0)) + scale_y_continuous(expand = c(0,0)) + coord_cartesian(xlim = c(0, 2000))+ # Adjust the cutoff point here theme(panel.background = element_blank(), # Make the background clear axis.line = element_line(color = "black"))
解答
geom_smooth()并不适用于你的场景——它的作用是对原始观测数据做平滑拟合(比如LOESS、线性回归曲线),而你现在是在绘制预先计算好的理论正态分布曲线,所以用它不对路。
要让你的正态曲线更平滑,有两种简单方法:
方法1:调整stat_function的采样点数
stat_function默认只使用100个点来绘制曲线,点数少就会显得不够平滑。只需给它增加n参数,设置更多采样点即可,比如n=1000:
# 修改后的stat_function部分 stat_function(fun = function(x) dnorm(x, mean = mean_m1, sd = sd_m1) * bin_width * total_points_m1, col= '#8B0000', n = 1000) + stat_function(fun = function(x) dnorm(x, mean = mean_m2, sd = sd_m2) * bin_width * total_points_m2, col= '#8B0000', n = 1000) +
方法2:提前生成曲线数据框,用geom_line绘制
这种方式更灵活,你可以完全控制曲线的取值范围和点数:
# 先生成两条正态曲线的数据集 curve_m1 <- data.frame( x = seq(min(data_m1$masses_kDa), max(data_m1$masses_kDa), length.out = 1000), y = dnorm(seq(min(data_m1$masses_kDa), max(data_m1$masses_kDa), length.out = 1000), mean = mean_m1, sd = sd_m1) * bin_width * total_points_m1 ) curve_m2 <- data.frame( x = seq(min(data_m2$masses_kDa), max(data_m2$masses_kDa), length.out = 1000), y = dnorm(seq(min(data_m2$masses_kDa), max(data_m2$masses_kDa), length.out = 1000), mean = mean_m2, sd = sd_m2) * bin_width * total_points_m2 ) # 绘图时替换stat_function为geom_line ggplot(cleaned_data, aes(x=masses_kDa)) + geom_histogram(color='transparent', fill='red', alpha = 0.5, bins= 50) + labs(x=' ', y=' ', title=' ') + geom_line(data = curve_m1, aes(x = x, y = y), col= '#8B0000') + geom_line(data = curve_m2, aes(x = x, y = y), col= '#8B0000') + geom_vline(aes(xintercept = mean_m1), linetype="dashed", color = 'black') + geom_vline(aes(xintercept = mean_m2), linetype="dashed", color = 'black') + scale_x_continuous(expand = c(0,0)) + scale_y_continuous(expand = c(0,0)) + coord_cartesian(xlim = c(0, 2000))+ theme(panel.background = element_blank(), axis.line = element_line(color = "black"))
两种方法都能让曲线变得更平滑,方法1更简洁,方法2适合需要对曲线做更多自定义调整的场景。
内容的提问来源于stack exchange,提问作者user124413
相关产品推荐
相关产品推荐

