如何在R语言中基于样本量合并动态生成的分组?
Great question! When dealing with dynamically created bins that have uneven frequencies (like your tiny last two bins), merging small groups based on sample size is a common task. Here's how you can do it in R, step by step:
Step 1: Prepare Your Frequency Data
First, let's formalize the frequency table from your data and extract the upper/lower bounds of each bin (this makes merging labels cleaner later):
# Your original dynamic bin creation code (adapts to your data's min/max) min_y <- min(sr$y) max_y <- max(sr$y) groups <- seq(min_y, max_y, (max_y - min_y)/10) # Generate frequency table from your sample data freq_table <- table(cut(sr$y, breaks = groups, include.lowest = TRUE)) # Convert to a data frame for easier manipulation freq_df <- data.frame( group = names(freq_table), count = as.numeric(freq_table), stringsAsFactors = FALSE ) # Extract lower and upper bounds from bin labels (works for default cut formatting) freq_df$lower <- as.numeric(gsub("\\[(\\d+),.*", "\\1", freq_df$group)) freq_df$upper <- as.numeric(gsub(".*,(\\d+)\\]", "\\1", freq_df$group))
Step 2: Define a Threshold and Merge Small Bins
Next, set a minimum sample size threshold (adjust this to your needs—e.g., 500 in this example) and iterate through the bins to merge small adjacent groups:
# Set your minimum desired sample size per merged bin min_count <- 500 # Initialize a list to hold merged bins merged_bins <- list() # Start with the first bin current_bin <- list( lower = freq_df$lower[1], upper = freq_df$upper[1], count = freq_df$count[1] ) # Iterate through remaining bins to merge small groups for (i in 2:nrow(freq_df)) { if (current_bin$count < min_count) { # Merge current bin with the next one current_bin$upper <- freq_df$upper[i] current_bin$count <- current_bin$count + freq_df$count[i] } else { # Add the valid bin to our results and start fresh with the next bin merged_bins[[length(merged_bins) + 1]] <- current_bin current_bin <- list( lower = freq_df$lower[i], upper = freq_df$upper[i], count = freq_df$count[i] ) } } # Don't forget to add the last bin (even if it's below the threshold) merged_bins[[length(merged_bins) + 1]] <- current_bin # Convert merged bins to a clean, readable data frame merged_df <- do.call(rbind, merged_bins) merged_df <- as.data.frame(merged_df) merged_df$group_label <- paste(merged_df$lower, merged_df$upper, sep = "-") # View the final merged table print(merged_df[, c("group_label", "count")])
How It Works
- We start with the first bin and keep merging it with subsequent bins until its total sample size meets or exceeds your threshold.
- If a bin already meets the threshold, we add it to our results and move to the next bin.
- The final bin is always added, even if it's below the threshold (you can adjust this logic if you want to force merging it with the previous bin instead).
Example Output for Your Data
Using your sample frequencies and min_count = 500, the merged table would look like this:
| group_label | count |
|---|---|
| 0-10 | 1700 |
| 10-15 | 1400 |
| 15-20 | 1700 |
| 20-25 | 1900 |
| 25-30 | 1500 |
| 30-35 | 1000 |
| 35-50 | 351 |
The last three small bins (35-40, 40-45, 45-50) are merged into one since their combined count (300+50+1=351) is still below the threshold, but we can't merge further left without breaking the threshold for the 30-35 bin.
Adjustments You Can Make
- Change the threshold: Modify
min_countto match your needs (e.g.,1000would merge the 0-5 and 5-10 bins into one). - Merge from right to left: Reverse the
freq_dfbefore iterating if you prefer to merge small bins starting from the upper end. - Force final bin merge: If you don't want any bins below the threshold, add logic to merge the final bin with the previous one if it's too small.
内容的提问来源于stack exchange,提问作者Aldi Kasse 2

