基于名称匹配的行间T检验:样本与参考组的p值提取方法
Got it, let's break this down step by step to get those p-values you need. First, let's make sure we're working with the same data, then we'll cover two approaches—one with base R, and another using the tidyverse for cleaner, more reproducible output.
Step 1: Load Your Data
First, let's load the data you provided so we can work with it directly:
# Load the dataset data <- structure(list(Area = c(3.62, 12.14, 13.74, 3.14, 131.86, 7.14, 14.66, 15.38, 12, 13.02, 192.12, 21.54, 13.38, 55.84, 23.86, 5.38, 200.56, 11.06, 12.12, 41.98, 33.32, 6.96, 210.12, 3.56), Names = c("Mark", "Peter", "Greg", "Manuel", "Reference1", "Reference2", "Mark", "Peter", "Greg", "Manuel", "Reference1", "Reference2", "Mark", "Peter", "Greg", "Manuel", "Reference1", "Reference2", "Mark", "Peter", "Greg", "Manuel", "Reference1", "Reference2")), row.names = c(NA, -24L), class = "data.frame")
Quick check to confirm each group has 4 observations (which matches your requirement):
# Verify observation counts per group table(data$Names)
Approach 1: Base R (No Extra Packages Needed)
If you prefer sticking to base R, this loop-based method will collect all p-values into a neat data frame:
# Extract the two reference groups' Area values ref1 <- data$Area[data$Names == "Reference1"] ref2 <- data$Area[data$Names == "Reference2"] # List of samples we need to test target_samples <- c("Mark", "Peter", "Greg", "Manuel") # Initialize an empty data frame to store results results_df <- data.frame( Sample = character(), Reference = character(), p_value = numeric(), stringsAsFactors = FALSE ) # Loop through each sample and run t-tests against both references for(sample in target_samples) { # Get the Area values for the current sample sample_vals <- data$Area[data$Names == sample] # Test against Reference1 and add to results test_ref1 <- t.test(sample_vals, ref1) results_df <- rbind(results_df, data.frame(Sample = sample, Reference = "Reference1", p_value = test_ref1$p.value)) # Test against Reference2 and add to results test_ref2 <- t.test(sample_vals, ref2) results_df <- rbind(results_df, data.frame(Sample = sample, Reference = "Reference2", p_value = test_ref2$p.value)) } # View the final results results_df
Approach 2: Tidyverse + Broom (Cleaner, Reproducible)
If you like working with tidy data principles, this method uses dplyr for data manipulation and broom to convert test results into a data frame:
First, install and load the required packages if you haven't already:
# Install packages (run once) if(!require(tidyverse)) install.packages("tidyverse") if(!require(broom)) install.packages("broom") # Load packages library(tidyverse) library(broom)
Now, process the data and run the tests:
# Split data into sample and reference subsets sample_subset <- data %>% filter(Names %in% c("Mark", "Peter", "Greg", "Manuel")) %>% rename(Sample = Names) ref_subset <- data %>% filter(Names %in% c("Reference1", "Reference2")) %>% rename(Reference = Names) # Create all possible sample-reference pairs test_combinations <- expand_grid( Sample = unique(sample_subset$Sample), Reference = unique(ref_subset$Reference) ) # Define a helper function to run t-test and extract p-value calculate_p <- function(sample_name, ref_name) { sample_data <- sample_subset$Area[sample_subset$Sample == sample_name] ref_data <- ref_subset$Area[ref_subset$Reference == ref_name] t_test_result <- t.test(sample_data, ref_data) return(t_test_result$p.value) } # Apply the function to all pairs final_results <- test_combinations %>% mutate(p_value = map2_dbl(Sample, Reference, calculate_p)) # View the results final_results
Notes on the T-Test
By default, t.test() uses the Welch two-sample t-test (which doesn't assume equal variances between groups). If you need to assume equal variances, add the var.equal = TRUE argument to the t.test() call, like this:
t.test(sample_vals, ref1, var.equal = TRUE)
内容的提问来源于stack exchange,提问作者Rechlay

