如何在R中对期中成绩及分数差值执行双尾两样本t检验
Got it, let's work through your two-sample t-test problem—since you only have summary statistics (not raw student scores) for the S and L cohorts, we need to use methods tailored for aggregated data instead of R's default t.test() which expects individual observations. Here's a step-by-step breakdown:
Your original read.table code has formatting issues (missing line breaks and incomplete data for 2017_S). Let's correct that first—fill in the missing difference stats with placeholder values (replace these with your actual data!):
d <- read.table(text = "Cohort N Midterm_Mean Midterm_SD Final_Mean Final_SD Diff_Mean Diff_SD 2016_L 38 77.4 3.0 73.7 4.2 -3.7 2.1 2017_S 37 81.9 2.1 70.0 4.6 -11.9 3.2", # Replace Diff_Mean/Diff_SD with your real values header = TRUE, stringsAsFactors = FALSE)
We'll cover two approaches: using a dedicated package (the easiest way) and manual calculation (for full understanding).
Option 1: Use the BSDA Package (Simplest)
The BSDA library has a tsum.test() function built specifically for t-tests with summary stats. First install and load it:
install.packages("BSDA") # Run once only library(BSDA)
Extract the midterm stats for each cohort:
# 2016_L cohort stats n_L <- d$N[d$Cohort == "2016_L"] mean_mid_L <- d$Midterm_Mean[d$Cohort == "2016_L"] sd_mid_L <- d$Midterm_SD[d$Cohort == "2016_L"] # 2017_S cohort stats n_S <- d$N[d$Cohort == "2017_S"] mean_mid_S <- d$Midterm_Mean[d$Cohort == "2017_S"] sd_mid_S <- d$Midterm_SD[d$Cohort == "2017_S"]
Run the two-tailed t-tests:
# Welch t-test (for unequal variances, more conservative) midterm_welch <- tsum.test(mean.x = mean_mid_L, s.x = sd_mid_L, n.x = n_L, mean.y = mean_mid_S, s.y = sd_mid_S, n.y = n_S, alternative = "two.sided") # Pooled t-test (for equal variances) midterm_pooled <- tsum.test(mean.x = mean_mid_L, s.x = sd_mid_L, n.x = n_L, mean.y = mean_mid_S, s.y = sd_mid_S, n.y = n_S, alternative = "two.sided", var.equal = TRUE) # View results print(midterm_welch) print(midterm_pooled)
Option 2: Manual Calculation (For Transparency)
If you prefer not to use external packages, calculate the t-statistic and p-value by hand:
Welch T-Test (Unequal Variances)
# Calculate t-statistic t_welch <- (mean_mid_L - mean_mid_S) / sqrt( (sd_mid_L^2/n_L) + (sd_mid_S^2/n_S) ) # Calculate Welch-Satterthwaite degrees of freedom df_welch <- ( (sd_mid_L^2/n_L + sd_mid_S^2/n_S)^2 ) / ( (sd_mid_L^2/n_L)^2/(n_L-1) + (sd_mid_S^2/n_S)^2/(n_S-1) ) # Two-tailed p-value p_welch <- 2 * pt(abs(t_welch), df = df_welch, lower.tail = FALSE) # Print results cat("Welch Two-Sample t-test for Midterm Scores:\n") cat(sprintf("t = %.3f, df = %.2f, p-value = %.4f\n", t_welch, df_welch, p_welch))
Pooled T-Test (Equal Variances)
# Calculate pooled standard deviation sp <- sqrt( ( (n_L-1)*sd_mid_L^2 + (n_S-1)*sd_mid_S^2 ) / (n_L + n_S - 2) ) # Calculate t-statistic t_pooled <- (mean_mid_L - mean_mid_S) / (sp * sqrt(1/n_L + 1/n_S)) # Degrees of freedom df_pooled <- n_L + n_S - 2 # Two-tailed p-value p_pooled <- 2 * pt(abs(t_pooled), df = df_pooled, lower.tail = FALSE) # Print results cat("\nPooled Two-Sample t-test for Midterm Scores:\n") cat(sprintf("t = %.3f, df = %.0f, p-value = %.4f\n", t_pooled, df_pooled, p_pooled))
Repeat the same process, but use the difference stats (Diff_Mean, Diff_SD) instead of midterm scores.
Using the BSDA Package
# Extract difference stats mean_diff_L <- d$Diff_Mean[d$Cohort == "2016_L"] sd_diff_L <- d$Diff_SD[d$Cohort == "2016_L"] mean_diff_S <- d$Diff_Mean[d$Cohort == "2017_S"] sd_diff_S <- d$Diff_SD[d$Cohort == "2017_S"] # Run tests diff_welch <- tsum.test(mean.x = mean_diff_L, s.x = sd_diff_L, n.x = n_L, mean.y = mean_diff_S, s.y = sd_diff_S, n.y = n_S, alternative = "two.sided") diff_pooled <- tsum.test(mean.x = mean_diff_L, s.x = sd_diff_L, n.x = n_L, mean.y = mean_diff_S, s.y = sd_diff_S, n.y = n_S, alternative = "two.sided", var.equal = TRUE) # View results print(diff_welch) print(diff_pooled)
Manual Calculation for Differences
Swap in the difference stats into the manual code from Section 2—here's a condensed version:
# Welch t-test for differences t_diff_welch <- (mean_diff_L - mean_diff_S) / sqrt( (sd_diff_L^2/n_L) + (sd_diff_S^2/n_S) ) df_diff_welch <- ( (sd_diff_L^2/n_L + sd_diff_S^2/n_S)^2 ) / ( (sd_diff_L^2/n_L)^2/(n_L-1) + (sd_diff_S^2/n_S)^2/(n_S-1) ) p_diff_welch <- 2 * pt(abs(t_diff_welch), df = df_diff_welch, lower.tail = FALSE) # Pooled t-test for differences sp_diff <- sqrt( ( (n_L-1)*sd_diff_L^2 + (n_S-1)*sd_diff_S^2 ) / (n_L + n_S - 2) ) t_diff_pooled <- (mean_diff_L - mean_diff_S) / (sp_diff * sqrt(1/n_L + 1/n_S)) df_diff_pooled <- n_L + n_S - 2 p_diff_pooled <- 2 * pt(abs(t_diff_pooled), df = df_diff_pooled, lower.tail = FALSE) # Print results cat("Welch Two-Sample t-test for Score Differences:\n") cat(sprintf("t = %.3f, df = %.2f, p-value = %.4f\n", t_diff_welch, df_diff_welch, p_diff_welch)) cat("\nPooled Two-Sample t-test for Score Differences:\n") cat(sprintf("t = %.3f, df = %.0f, p-value = %.4f\n", t_diff_pooled, df_diff_pooled, p_diff_pooled))
- Variance Assumption: If the two cohorts' standard deviations are more than 2x apart, stick with the Welch test. If they're similar, the pooled test is acceptable.
- Placeholder Data: Remember to replace the
Diff_MeanandDiff_SDvalues for 2017_S with your actual data—my placeholders are just for demonstration.
内容的提问来源于stack exchange,提问作者mojo

