基于for循环的蒙特卡洛模拟优化中采样过程分数最大值的监控实现与代码改进问询
Great question—let's tackle this step by step. The core issue with your initial loop is that you're trying to add a new Score_max column inside the rbind call, which creates mismatched column counts on subsequent iterations. Here are two clean, base-R solutions that avoid this problem and efficiently track the cumulative maximum score, just like you need:
Solution 1: Track Maximum in Real Time (Most Efficient for Large n)
This approach avoids repeated rbind calls (which are slow for large datasets) by maintaining a running maximum value as you sample. We pre-allocate storage vectors for speed, then update the maximum with each iteration.
# Prepare your base data and sampling function (unchanged from your example) df <- expand.grid(A = seq(0,1,length.out=20), B = seq(0,2,length.out=20), C = seq(0,3,length.out=20)) %>% distinct() sampling_function <- function(x){ repeat { Point <- x[sample(nrow(x), size=1), 1:3] Point$Score <- Point$A * Point$B * Point$C if(Point$Score >= 1) break } return(Point) } # Initialize storage for results n <- 1000 sampling_numbers <- 1:n cumulative_max <- numeric(n) current_max <- -Inf # Start with a value lower than any possible score # Run sampling and track maximum for(k in 1:n){ sampled_point <- sampling_function(df) current_max <- max(current_max, sampled_point$Score) cumulative_max[k] <- current_max } # Build your target data frame result_df <- data.frame( Sampling_number = sampling_numbers, Score_max = cumulative_max ) # Preview the first few rows head(result_df)
Solution 2: Collect All Scores First, Then Compute Cumulative Max
If you prefer to first capture all sampled scores before calculating the running maximum, base R's cummax() function is exactly what you're asking for—it generates a vector where each value is the maximum of all previous values up to that point.
# Reuse your data and sampling function df <- expand.grid(A = seq(0,1,length.out=20), B = seq(0,2,length.out=20), C = seq(0,3,length.out=20)) %>% distinct() sampling_function <- function(x){ repeat { Point <- x[sample(nrow(x), size=1), 1:3] Point$Score <- Point$A * Point$B * Point$C if(Point$Score >= 1) break } return(Point) } n <- 1000 # Pre-allocate data frame to store all samples all_samples <- data.frame( Sampling_number = 1:n, Score = numeric(n) ) # Fill in scores with your sampling function for(k in 1:n){ all_samples$Score[k] <- sampling_function(df)$Score } # Calculate cumulative maximum with cummax() all_samples$Score_max <- cummax(all_samples$Score) # Preview results head(all_samples)
Visualize the Cumulative Max
To see when your score stabilizes (and determine how many samples you need), use base R's plotting tools:
plot(result_df$Sampling_number, result_df$Score_max, type = "l", lwd = 2, col = "steelblue", xlab = "Sampling Number", ylab = "Cumulative Maximum Score", main = "Monte Carlo Optimization: Maximum Score Over Samples")
Why Your Initial Code Failed
Your original loop tried to add the Score_max column mid-iteration: after the first loop, m.c had 5 columns (A, B, C, Score, Score_max), but data.frame(P) only had 4 columns. This mismatched column count caused the rbind error. Both solutions above avoid this by either:
- Tracking the maximum outside the data frame first, or
- Adding the cumulative max column after all samples are collected.
内容的提问来源于stack exchange,提问作者Marcel

