You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Bootstrap法计算率比的P值以检验其显著性?

Hey there! Let's break down how to calculate a Bootstrap p-value to test your rate ratio's significance—this is totally doable, and I’ll walk you through the key steps with a conceptual code example to make it concrete.

Bootstrap P-Value for Rate Ratio Significance Testing

Core Concept

The null hypothesis we’re testing here is that the true rate ratio equals 1 (meaning no difference between your two groups). Unlike how you calculated the confidence interval (resampling groups separately), for the p-value we need to generate a distribution of rate ratios that would exist if the null hypothesis were true, then compare your observed rate ratio to that distribution.

Step-by-Step Process

1. Set Up Null Hypothesis Resampling

Instead of resampling each group independently (which reflects the observed data's sampling distribution), we first pool all your data to simulate the scenario where group membership doesn’t affect the rate:

  • Combine all events and person-time (or your rate denominator) from both groups into a single pool.
  • For each Bootstrap iteration:
    • Resample with replacement from this pooled dataset, making sure you keep the original group sizes (e.g., if Group A had 50 observations and Group B had 60, each resample should have 50 "A" and 60 "B" entries drawn from the pool).
    • Calculate the rate ratio for this resampled dataset.

2. Generate the Null Bootstrap Distribution

Run this resampling loop at least 10,000 times (more iterations = more stable results). This will give you a distribution of rate ratios that would occur by chance if the null hypothesis were true.

3. Calculate the P-Value

There are two standard approaches, depending on your hypothesis:

  • Two-tailed test (most common):
    1. Take your observed rate ratio (RR_obs).
    2. Count how many Bootstrap rate ratios are either ≤ 1/RR_obs or ≥ RR_obs (since rate ratios are multiplicative, the null is symmetric around 1).
    3. Divide that count by the total number of iterations—this is your two-tailed p-value.
  • One-tailed test (directional hypothesis):
    1. If you expect RR_obs > 1, count how many Bootstrap rate ratios are ≥ RR_obs.
    2. If you expect RR_obs < 1, count how many are ≤ RR_obs.
    3. Divide by total iterations for your one-tailed p-value.

Conceptual Code Example (R)

Here’s a rough sketch to illustrate how this might work in practice (adjust based on your data structure):

# Assume you have aggregated data for two groups
events_a <- 15  # Number of events in Group A
person_time_a <- 100  # Total person-time in Group A
events_b <- 8   # Number of events in Group B
person_time_b <- 100  # Total person-time in Group B

# Calculate your observed rate ratio
observed_rr <- (events_a / person_time_a) / (events_b / person_time_b)

# For null resampling, we'll work with individual-level data (expand counts)
# Create a vector of 1s (events) and 0s (non-events) for both groups
group_a_data <- c(rep(1, events_a), rep(0, person_time_a - events_a))
group_b_data <- c(rep(1, events_b), rep(0, person_time_b - events_b))

# Pool all data
pooled_data <- c(group_a_data, group_b_data)
group_labels <- c(rep("A", length(group_a_data)), rep("B", length(group_b_data)))

# Function to generate one null Bootstrap rate ratio
generate_null_rr <- function() {
  # Resample while preserving original group sizes
  resampled_idx <- sample(seq_along(pooled_data), size = length(pooled_data), replace = TRUE)
  resampled_events <- pooled_data[resampled_idx]
  resampled_groups <- group_labels[resampled_idx]
  
  # Calculate rates for resampled groups
  rate_a <- sum(resampled_events[resampled_groups == "A"]) / length(resampled_events[resampled_groups == "A"])
  rate_b <- sum(resampled_events[resampled_groups == "B"]) / length(resampled_events[resampled_groups == "B"])
  
  # Avoid division by zero (add a tiny value if needed)
  if (rate_b == 0) return(Inf)
  return(rate_a / rate_b)
}

# Run 10,000 Bootstrap iterations
null_rr_dist <- replicate(10000, generate_null_rr())

# Calculate two-tailed p-value
lower_tail_count <- sum(null_rr_dist <= 1 / observed_rr)
upper_tail_count <- sum(null_rr_dist >= observed_rr)
p_value <- (lower_tail_count + upper_tail_count) / 10000

cat("Two-tailed Bootstrap p-value:", round(p_value, 4), "\n")

Quick Tips

  • Don’t skip pooling: Resampling groups separately (like for CI) gives you the sampling distribution of the observed rate ratio, not the null distribution—this is a common mistake!
  • Handle edge cases: If you have zero events in a resampled group, add a small epsilon (like 0.001) to the denominator to avoid division by zero.
  • Iteration count: 10,000 is a good minimum; if you need more precision, bump it up to 20,000 or 50,000.

内容的提问来源于stack exchange,提问作者AprilCS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:23:08