在R语言中绘制多连续变量频率直方图并保持统一刻度范围
Hey there! I totally get your frustration with multi.hist not letting you set uniform scales—it makes comparing distributions across variables way harder. Let's fix this with ggplot2 (even if you're new to R, I'll break it down super clearly), and we'll also incorporate that topic variable to make the visualization more insightful.
First, let's start with some setup using a simulated dataset that matches your description (you can swap this out for your real Data):
# 模拟你的数据(替换成你真实的Data即可) set.seed(123) # 保证结果可重复 Data <- data.frame( topic = factor(sample(1:5, 100, replace = TRUE)), # 分类变量:主题1-5 var1 = rnorm(100, mean = 5, sd = 2), # 连续变量1 var2 = rnorm(100, mean = 7, sd = 1.5), # 连续变量2 var3 = rnorm(100, mean = 4, sd = 2.5) # 连续变量3 )
第一步:安装并加载必要的包
If you haven't installed these yet, run this first:
install.packages(c("ggplot2", "tidyr"))
Then load the packages:
library(ggplot2) library(tidyr)
第二步:转换数据格式(宽→长)
ggplot2 works best with long-format data (each row represents one value from one variable). We'll use pivot_longer to convert your wide data:
long_data <- pivot_longer( Data, cols = starts_with("var"), # 替换成你所有连续变量的列名,比如c("age", "income", "score") names_to = "variable", # 新列名:存储原来的变量名称 values_to = "value" # 新列名:存储对应变量的数值 )
第三步:绘制统一刻度的基础直方图(不含分类变量)
The magic here is scales = "fixed"—it forces all facets to use the exact same x and y axis scales, which solves your multi.hist problem:
ggplot(long_data, aes(x = value)) + geom_histogram(bins = 10, fill = "steelblue", color = "white") + # 调整bins可以改变直方图的组数 facet_wrap(~variable, scales = "fixed") + # 按变量分面,统一刻度 labs( title = "Distribution of Continuous Variables (Uniform Scales)", x = "Value", y = "Count" ) + theme_minimal() # 简洁美观的主题,也可以用theme_bw()
第四步:加入分类变量(主题1-5)的对比
We can integrate the topic variable in two useful ways—pick whichever fits your analysis needs:
方式1:每个变量的直方图按主题并列显示
ggplot(long_data, aes(x = value, fill = topic)) + geom_histogram( bins = 10, position = "dodge", # 并列显示不同主题的直方图;换成"stack"就是堆叠样式 color = "white", alpha = 0.8 # 增加透明度,避免颜色重叠看不清 ) + facet_wrap(~variable, scales = "fixed") + labs( title = "Variable Distributions by Topic", x = "Value", y = "Count", fill = "Topic" ) + theme_minimal()
方式2:按主题分面,每个主题下对比所有变量
ggplot(long_data, aes(x = value, fill = variable)) + geom_histogram( bins = 10, position = "dodge", color = "white", alpha = 0.8 ) + facet_wrap(~topic, scales = "fixed") + labs( title = "Variable Distributions Within Each Topic", x = "Value", y = "Count", fill = "Variable" ) + theme_minimal()
可选:用base R的multi.hist实现统一刻度
If you still prefer using multi.hist, you can manually calculate global scales and apply them:
library(psych) # 提取所有连续变量,计算全局范围和最大频数 continuous_vars <- Data[, c("var1", "var2", "var3")] # 替换成你的连续变量列 all_values <- unlist(continuous_vars) x_range <- range(all_values) max_count <- max(sapply(continuous_vars, function(x) hist(x, plot = FALSE)$counts)) # 绘制统一刻度的多直方图 multi.hist( continuous_vars, xlim = x_range, ylim = c(0, max_count), col = "steelblue" )
内容的提问来源于stack exchange,提问作者Henrik

