如何用purrr的map函数逐行执行prop.test并将结果添加到数据框?
Hey there! I’ve dealt with this exact problem before, so let me walk you through two straightforward approaches to get those p-values and confidence intervals (CI) added to your dataframe. We’ll cover both base R and tidyverse methods so you can pick what fits your workflow best.
First, Let’s Create Sample Data
To make this concrete, let’s start with a dummy dataframe that matches your structure (successes + total trials):
set.seed(123) # Ensures reproducible results df <- data.frame( successes = sample(1:50, 10), trials = sample(50:100, 10) )
Approach 1: Base R with apply()
This method uses base R tools, no extra packages needed. We’ll write a custom function to process each row, then merge the results back into your original dataframe.
Step 1: Define a Row-Processing Function
This function takes a row from your dataframe, runs the proportion test, and extracts the metrics you care about:
process_row <- function(row) { # Pull success count and total trials from the row success_num <- row["successes"] total_trials <- row["trials"] # Run prop.test (swap with binom.test if you need exact binomial tests) test_output <- prop.test(x = success_num, n = total_trials, correct = FALSE) # Note: `correct=FALSE` turns off continuity correction; adjust based on your needs # Extract key metrics p_val <- test_output$p.value ci_lower <- test_output$conf.int[1] ci_upper <- test_output$conf.int[2] # Format CI as a readable string (e.g., "(0.32, 0.58)") ci_string <- sprintf("(%.2f, %.2f)", ci_lower, ci_upper) # Return results as a named vector return(c(p_value = p_val, ci = ci_string, ci_low = ci_lower, ci_high = ci_upper)) }
Step 2: Apply the Function and Merge Results
# Apply the function to every row, transpose the output to a dataframe test_results <- as.data.frame(t(apply(df, 1, process_row))) # Merge with original dataframe final_df <- cbind(df, test_results) # View the result head(final_df)
Approach 2: Tidyverse with dplyr + broom
If you prefer the tidyverse workflow, using dplyr::rowwise() and the broom package (which converts statistical outputs to tidy data frames) makes this super clean.
Step 1: Install/Load Required Packages
install.packages(c("dplyr", "broom")) # Run once if not installed library(dplyr) library(broom)
Step 2: Process Rows and Extract Results
final_df_tidy <- df %>% rowwise() %>% # Treat each row as an individual group mutate( # Run the test and store the result as a list column test_result = list(prop.test(x = successes, n = trials, correct = FALSE)), # Extract p-value using broom::tidy() p_value = tidy(test_result)$p.value, # Extract CI bounds ci_low = tidy(test_result)$conf.low, ci_high = tidy(test_result)$conf.high, # Format CI string ci = sprintf("(%.2f, %.2f)", ci_low, ci_high) ) %>% ungroup() %>% # Reset row-wise grouping select(-test_result) # Optional: Remove the raw test result column # View the tidy result head(final_df_tidy)
Key Notes
- Switching to
binom.test: Just replaceprop.testwithbinom.testin either method—both functions return objects with the same structure for p-values and confidence intervals. - Continuity Correction: The
correct=FALSEargument inprop.testturns off continuity correction. If you want it enabled (the default), just remove that argument. - Additional Metrics: If you want to include the estimated proportion, add
estimate = tidy(test_result)$estimate(tidyverse) orestimate = test_output$estimate(base R) to your output.
内容的提问来源于stack exchange,提问作者Luise

