You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中基于样本量合并动态生成的分组?

Merging Bins by Sample Size in R

Great question! When dealing with dynamically created bins that have uneven frequencies (like your tiny last two bins), merging small groups based on sample size is a common task. Here's how you can do it in R, step by step:

Step 1: Prepare Your Frequency Data

First, let's formalize the frequency table from your data and extract the upper/lower bounds of each bin (this makes merging labels cleaner later):

# Your original dynamic bin creation code (adapts to your data's min/max)
min_y <- min(sr$y)
max_y <- max(sr$y)
groups <- seq(min_y, max_y, (max_y - min_y)/10)

# Generate frequency table from your sample data
freq_table <- table(cut(sr$y, breaks = groups, include.lowest = TRUE))

# Convert to a data frame for easier manipulation
freq_df <- data.frame(
  group = names(freq_table),
  count = as.numeric(freq_table),
  stringsAsFactors = FALSE
)

# Extract lower and upper bounds from bin labels (works for default cut formatting)
freq_df$lower <- as.numeric(gsub("\\[(\\d+),.*", "\\1", freq_df$group))
freq_df$upper <- as.numeric(gsub(".*,(\\d+)\\]", "\\1", freq_df$group))

Step 2: Define a Threshold and Merge Small Bins

Next, set a minimum sample size threshold (adjust this to your needs—e.g., 500 in this example) and iterate through the bins to merge small adjacent groups:

# Set your minimum desired sample size per merged bin
min_count <- 500

# Initialize a list to hold merged bins
merged_bins <- list()

# Start with the first bin
current_bin <- list(
  lower = freq_df$lower[1],
  upper = freq_df$upper[1],
  count = freq_df$count[1]
)

# Iterate through remaining bins to merge small groups
for (i in 2:nrow(freq_df)) {
  if (current_bin$count < min_count) {
    # Merge current bin with the next one
    current_bin$upper <- freq_df$upper[i]
    current_bin$count <- current_bin$count + freq_df$count[i]
  } else {
    # Add the valid bin to our results and start fresh with the next bin
    merged_bins[[length(merged_bins) + 1]] <- current_bin
    current_bin <- list(
      lower = freq_df$lower[i],
      upper = freq_df$upper[i],
      count = freq_df$count[i]
    )
  }
}

# Don't forget to add the last bin (even if it's below the threshold)
merged_bins[[length(merged_bins) + 1]] <- current_bin

# Convert merged bins to a clean, readable data frame
merged_df <- do.call(rbind, merged_bins)
merged_df <- as.data.frame(merged_df)
merged_df$group_label <- paste(merged_df$lower, merged_df$upper, sep = "-")

# View the final merged table
print(merged_df[, c("group_label", "count")])

How It Works

  • We start with the first bin and keep merging it with subsequent bins until its total sample size meets or exceeds your threshold.
  • If a bin already meets the threshold, we add it to our results and move to the next bin.
  • The final bin is always added, even if it's below the threshold (you can adjust this logic if you want to force merging it with the previous bin instead).

Example Output for Your Data

Using your sample frequencies and min_count = 500, the merged table would look like this:

group_labelcount
0-101700
10-151400
15-201700
20-251900
25-301500
30-351000
35-50351

The last three small bins (35-40, 40-45, 45-50) are merged into one since their combined count (300+50+1=351) is still below the threshold, but we can't merge further left without breaking the threshold for the 30-35 bin.

Adjustments You Can Make

  • Change the threshold: Modify min_count to match your needs (e.g., 1000 would merge the 0-5 and 5-10 bins into one).
  • Merge from right to left: Reverse the freq_df before iterating if you prefer to merge small bins starting from the upper end.
  • Force final bin merge: If you don't want any bins below the threshold, add logic to merge the final bin with the previous one if it's too small.

内容的提问来源于stack exchange,提问作者Aldi Kasse 2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:30:35