如何创建Y轴含多组同类别数值的堆叠条形图?已尝试未解决
Hey there! I get it—trying to make a stacked bar chart with extreme value differences can be tricky. Let's break this down step by step to get your plot working properly.
First, let's formalize your sample data into a usable R data frame (I filled in the truncated Probability_n4 value with reasonable placeholders for completeness):
library(tidyverse) # Create your sample data frame df <- tibble( SampleNumber = 1:5, Probability_n1 = c(1, 1, 1, 1, 1), Probability_n2 = c(1.30066666666667e-14, 3.375e-17, 3.698e-16, 2.03733333333333e-17, 2.16866666666667e-19), Probability_n3 = c(1.243e-28, 5.59956666666667e-35, 4.75333333333333e-34, 1.0569e-34, 1.21333333333333e-37), Probability_n4 = c(2.266e-43, 4.36333333333333e-49, 1.2e-48, 5.6e-49, 3.1e-51) )
The Core Problem: Data Format
Stacked bar charts in ggplot2 work best with long-format data (one row per observation-group combination), but your data is in wide format (one column per probability group). Let's convert it first:
# Convert to long format df_long <- df %>% pivot_longer( cols = starts_with("Probability_n"), names_to = "Probability_Group", values_to = "Probability" )
Creating the Stacked Bar Chart
Now we can build the basic stacked bar chart. But wait—your probability values have massive differences (from 1 down to 1e-51). If we plot them on a linear scale, all the smaller values will be completely hidden under the Probability_n1 bars. Here are a few solutions:
Option 1: Linear Scale (For Demonstration)
You can make the plot, but the smaller groups won't be visible at all:
ggplot(df_long, aes(x = factor(SampleNumber), y = Probability, fill = Probability_Group)) + geom_bar(stat = "identity") + labs( x = "Sample Number", y = "Probability", title = "Stacked Bar Chart of Probabilities (Linear Scale)", fill = "Probability Group" ) + theme_minimal()
Option 2: Logarithmic Scale (To See Small Values)
To visualize all groups, we can use a logarithmic y-scale. We'll add a tiny offset to avoid issues with zero values (a safe practice even if your data doesn't have zeros):
ggplot(df_long, aes(x = factor(SampleNumber), y = Probability + 1e-52, fill = Probability_Group)) + geom_bar(stat = "identity") + scale_y_log10(labels = scales::scientific) + labs( x = "Sample Number", y = "Probability (Log Scale)", title = "Stacked Bar Chart of Probabilities (Logarithmic Scale)", fill = "Probability Group" ) + theme_minimal()
Option 3: Faceted Bar Charts (Alternative Clear View)
If stacking still feels clunky, you can split each probability group into its own panel with facets—this makes each group's values easy to compare:
ggplot(df_long, aes(x = factor(SampleNumber), y = Probability, fill = Probability_Group)) + geom_bar(stat = "identity") + scale_y_log10(labels = scales::scientific) + facet_wrap(~Probability_Group, scales = "free_y") + labs( x = "Sample Number", y = "Probability", title = "Probabilities by Sample and Group (Faceted View)", fill = "Probability Group" ) + theme_minimal()
Key Takeaways
- Always convert wide data to long format for stacked plots in
ggplot2 - Extreme value differences require scale adjustments (like log scale) or alternative views (like facets) to make all groups visible
内容的提问来源于stack exchange,提问作者Omer Saleem

