R语言多列分布直方图坐标轴交换与数值显示修正问询
Hey there! Let's fix those two histogram issues you're dealing with for your large dataset. I'll walk you through practical solutions tailored to your existing R code, covering both base R and ggplot2 approaches.
问题拆解
You've got two key pain points right now:
- Need to swap the x-axis (variable values) and y-axis (frequency counts) to create horizontal histograms
- Want to replace scientific notation on axes with readable standard numeric formatting
方案1:修改Base R循环代码
If you prefer sticking with the base R workflow you started, here's how to adjust it:
1.1 全局关闭科学计数法
First, set a global option to force R to use standard numeric formatting instead of scientific notation:
options(scipen = 999) # Higher values mean stronger preference for non-scientific notation
1.2 绘制横向直方图(轴交换)
Base R's hist function doesn't directly support horizontal plots, but we can extract its statistical results and use barplot to create the horizontal version. Update your loop like this:
var_to_plot = c("BASKETS_NZ","PIS","PIS_AP","PIS_DV","PIS_PL","PIS_SDV", "PIS_SHOPS","PIS_SR", "QUANTITY") par(mfrow=c(3,3)) options(scipen = 999) # Apply non-scientific notation setting for(i in var_to_plot){ # Get histogram stats without plotting hist_data <- hist(WKA_ohneJB[,i], plot = FALSE) # Create horizontal barplot with swapped axes barplot(hist_data$counts, names.arg = hist_data$mids, horiz = TRUE, xlab = "频数", ylab = i, main = "") }
plot = FALSEletshistcalculate data without drawing the default vertical plothoriz = TRUEflips the bars to horizontal, swapping the axis rolesnames.arg = hist_data$midsuses the original histogram's midpoints as y-axis labels
方案2:用ggplot2绘制更整洁的横向直方图
If you're open to using ggplot2 (you already tried it earlier), this method is more intuitive and flexible:
library(ggplot2) library(tidyr) # Reshape data to long format (cleaner than your original melt approach) df_long <- WKA_ohneJB %>% select(all_of(var_to_plot)) %>% pivot_longer(cols = everything(), names_to = "Variable", values_to = "Value") # Draw horizontal histograms with standard numeric formatting ggplot(df_long, aes(x = Value)) + geom_histogram(bins = 30, fill = "#2E86AB", color = "white") + facet_wrap(~Variable, ncol = 3) + # Arrange plots in 3 columns coord_flip() + # Swap x and y axes directly scale_x_continuous(labels = scales::comma) + # Use comma-separated numbers instead of scientific notation labs(x = "频数", y = "变量值") + theme_minimal()
coord_flip()instantly swaps the axes with no extra workscale_x_continuous(labels = scales::comma)ensures axes use readable standard formatting with thousands separatorsfacet_wrapautomatically organizes all variable plots into a neat grid
额外提示
For your 820k-row dataset:
- In ggplot2, adjust the
binsparameter ingeom_histogramto balance plot speed and clarity - In base R, use the
breaksargument inhistto control how data is grouped if needed
数据集片段参考
structure(list(X = c(1L, 2L, 3L, 4L, 5L, 6L, 7L, 8L, 9L, 10L, 821039L, 821040L, 821041L, 821042L, 821043L, 821044L, 821045L, 821046L, 821047L, 821048L), BASKETS_NZ = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L), LOGONS = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 0L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L), PIS = c(71L, 39L, 50L, 4L, 13L, 4L, 30L, 65L, 13L, 31L, 111L, 33L, 3L, 46L, 11L, 8L, 17L, 68L, 65L, 15L), PIS_AP = c(14L, 2L, 4L, 0L, 0L, 0L, 1L, 0L, 2L, 1L, 13L, 0L, 0L, 2L, 1L, 0L, 3L, 8L, 0L, 1L), PIS_DV = c(3L, 19L, 4L, 1L, 0L, 0L, 6L, 2L, 2L, 3L, 38L, 8L, 0L, 5L, 2L, 0L, 1L, 0L, 3L, 2L), PIS_PL = c(0L, 5L, 8L, 2L, 0L, 0L, 0L, 24L, 0L, 6L, 32L, 8L, 0L, 0L, 4L, 0L, 0L, 0L, 0L, 0L), PIS_SDV = c(18L, 0L, 11L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 6L, 0L, 0L, 13L, 0L, 0L, 1L, 15L, 1L, 0L), PIS_SHOPS = c(3L, 24L, 13L, 3L, 0L, 0L, 6L, 28L, 2L, 11L, 71L, 16L, 2L, 5L, 6L, 0L, 1L, 0L, 3L, 2L), PIS_SR = c(19L, 0L, 14L, 0L, 0L, 0L, 2L, 23L, 0L, 3L, 6L, 0L, 0L, 20L, 0L, 0L, 3L, 32L, 1L, 0L), QUANTITY = c(13L, 2L, 18L, 1L, 14L, 1L, 4L, 2L, 5L, 1L, 5L, 2L, 2L, 4L, 1L, 3L, 2L, 8L, 17L, 8L), WKA = c(1L, 1L, 1L, 1L, 1L, 1L, 0L, 0L, 1L, 0L, 1L, 1L, 1L, 1L, 1L, 1L, 0L, 0L, 1L, 1L), NEW_CUST = c(0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L), EXIST_CUST = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 0L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L), WEB_CUST = c(1L, 0L, 0L, 0L, 1L, 1L, 0L, 1L, 1L, 1L, 1L, 1L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 1L), MOBILE_CUST = c(0L, 1L, 1L, 1L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 0L, 1L, 0L), TABLET_CUST = c(0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 1L, 1L, 0L, 1L, 0L, 0L), LOGON_CUST_STEP2 = c(0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L)), row.names = c(1L, 2L, 3L, 4L, 5L, 6L, 7L, 8L, 9L, 10L, 821039L, 821040L, 821041L, 821042L, 821043L, 821044L, 821045L, 821046L, 821047L, 821048L ), class = "data.frame")
内容的提问来源于stack exchange,提问作者Kitty123

