如何以Tidy风格为数据框每行多列应用带独立参数的函数?
Hey there! Let's walk through how to solve this using tidyverse tools—no clunky for loops required. The goal is to apply a custom function that takes separate column values from each row as arguments, then populate the p_value1 and p_value2 columns with the results.
Step 1: Define Your Custom Function
First, let's formalize the function that will calculate the two p-values. I'll use common statistical tests as examples, but you can swap this out for your actual logic:
calculate_p_values <- function(success, trial, prog) { # Example 1: Binomial test p-value p1 <- binom.test(x = success, n = trial, p = prog)$p.value # Example 2: Poisson test p-value (alternative metric) p2 <- poisson.test(x = success, lambda = trial * prog)$p.value # Return both values as a named list list(p_value1 = p1, p_value2 = p2) }
Step 2: Apply the Function Row-Wise (Two Tidy Methods)
Method 1: Using dplyr::rowwise() (Most Intuitive)
This method is great for readability—we explicitly tell R to process the data row-by-row, then call our function with the relevant columns:
library(dplyr) # Process the data test_data_processed <- test_data %>% rowwise() %>% # Enable row-wise operations mutate( # Call our function with the three columns as arguments p_results = list(calculate_p_values(successes, trials, prognosis)), # Extract the two p-values into separate columns p_value1 = p_results$p_value1, p_value2 = p_results$p_value2 ) %>% select(-p_results) %>% # Remove the temporary list column ungroup() # Revert back to standard data frame behavior
Method 2: Using purrr::pmap() (Functional Programming Style)
If you prefer a more functional approach, pmap from the purrr package is perfect for mapping multiple arguments to a function across rows:
library(dplyr) library(purrr) test_data_processed <- test_data %>% mutate( # Map the three columns to our function's arguments p_results = pmap( .l = list(success = successes, trial = trials, prog = prognosis), .f = calculate_p_values ) ) %>% unnest_wider(p_results) # Expand the list column into separate columns
Key Notes
- Argument Matching: Make sure the order (or names) of the columns passed to the function match the function's parameter names. Using named lists in
pmap(likesuccess = successes) avoids mix-ups if your function has many parameters. - Ungrouping: After using
rowwise(), don't forget toungroup()if you plan on doing further column-wise operations—otherwise, R will keep treating the data as row-grouped. - Custom Logic: Replace the test calculations in
calculate_p_valueswith your actual function logic—this structure works for any function that takes multiple scalar inputs and returns multiple outputs.
Example Output Preview
After running either method, your test_data_processed will have populated p_value1 and p_value2 columns:
# A tibble: 10 × 5 successes trials prognosis p_value1 p_value2 <dbl> <dbl> <dbl> <dbl> <dbl> 1 0 50 0.0005 1.00e 0 1.00e 0 2 3 1788 0.001 9.77e-1 7.77e-1 3 6 1876 0.0025 2.60e-1 1.98e-1 # … with 7 more rows
内容的提问来源于stack exchange,提问作者Richi W

