在R中用ggplot绘制数据框多列折线图:波利亚urn模型可视化问题
Hey there! Let's sort out why your ggplot lines are getting all tangled up when plotting your Pólya urn simulation results. The core issue here is almost always not properly tracking which simulation run each data point belongs to when reshaping your data from wide to long format. Let's walk through the fix step by step.
First, Let's Make Sure Our Data is Structured Correctly
I'll start with reproducible code to generate the Pólya urn data matching your specs (1 white/1 black initial, 10 iterations, 5 simulations) — feel free to swap this out with your own test data if it's already generated:
# Set seed for consistent results set.seed(123) # Define simulation parameters num_simulations <- 5 num_iterations <- 10 # Initialize matrix to store white ball proportions urn_data_wide <- matrix(nrow = num_iterations, ncol = num_simulations) # Run the Pólya urn simulations for (sim in 1:num_simulations) { white <- 1 black <- 1 for (iter in 1:num_iterations) { # Draw a random ball draw <- sample(c("white", "black"), size = 1, prob = c(white, black)) # Add a matching color ball if (draw == "white") white <- white + 1 else black <- black + 1 # Store the proportion of white balls urn_data_wide[iter, sim] <- white / (white + black) } } # Convert to a wide-format data frame and add an iteration column urn_data_wide <- as.data.frame(urn_data_wide) colnames(urn_data_wide) <- paste0("sim_", 1:num_simulations) urn_data_wide$iteration <- 1:num_iterations
The Problem: Misconfigured Reshaping
If your lines are crossing randomly, you probably reshaped the data without explicitly linking each proportion to its simulation run and iteration. For example, a common mistake is using gather or pivot_longer without specifying which columns to reshape, leading to a messy long format that loses grouping context.
Correct Reshaping to Long Format
Use tidyr::pivot_longer to explicitly reshape only the simulation columns, keeping the iteration column as our reference point. This ensures we have a clear simulation group variable for ggplot to use:
library(tidyr) # Reshape to long format (this is the critical step!) urn_data_long <- pivot_longer( urn_data_wide, cols = starts_with("sim_"), # Target only the simulation columns names_to = "simulation", # Name for the new simulation group column values_to = "white_proportion" # Name for the proportion values )
Take a quick look at the long format data with head(urn_data_long) — you'll see each row has an iteration, simulation ID, and the corresponding white_proportion. This is the structure ggplot needs to draw separate lines per simulation.
Plotting with ggplot the Right Way
Now, when building your plot, you need to tell ggplot to group the data by simulation (using group = simulation) and map that group to color. This ensures each simulation's points are connected into a single, unbroken line:
library(ggplot2) ggplot(urn_data_long, aes(x = iteration, y = white_proportion)) + # Map simulation to both group and color for clear separation geom_line(aes(group = simulation, color = simulation), linewidth = 1) + geom_point(aes(color = simulation), size = 2) + # Add labels for clarity labs( title = "Pólya Urn Model: White Ball Proportion Over Iterations", x = "Iteration Number", y = "Proportion of White Balls", color = "Simulation Run" ) + theme_minimal()
Why This Works
By explicitly grouping the data by simulation, ggplot knows exactly which points belong to each run — no more random cross-connections between different simulations. The color mapping also makes it easy to distinguish each line at a glance.
内容的提问来源于stack exchange,提问作者Roni

