基于GLMM利用成绩数据判定不同题型的表现差异
Hey there! Let's walk through how to figure out if students perform better on Type A questions compared to Type B, using your 10 years of exam data. I'll break down the best analysis approaches based on your data structure, plus some extra considerations for that long-term dataset.
Your goal is to test if the average score on Type A is significantly higher than Type B. Since scores are continuous (0-100), we'll start with parametric tests (when assumptions hold) and use non-parametric alternatives if needed.
Scenario 1: Paired Data (Same Students Took Both Question Types)
If every student completed both Type A and Type B questions in each exam, your data is paired (two scores per student). This is the most common case for comparing two question types in the same test:
- First choice: Paired t-test
Requires that the difference between each student's A and B scores follows a normal distribution. Verify this with a Shapiro-Wilk test or QQ plot. - If normality fails: Wilcoxon Signed-Rank Test
A non-parametric alternative that doesn't rely on normal distribution assumptions.
Code Example (Python with Scipy)
from scipy.stats import ttest_rel, wilcoxon, shapiro # Assume a_scores and b_scores are arrays of corresponding student scores # First check normality of differences diff_scores = [a - b for a, b in zip(a_scores, b_scores)] shapiro_stat, shapiro_p = shapiro(diff_scores) print(f"Shapiro-Wilk normality test p-value: {shapiro_p:.4f}") # Paired t-test t_stat, t_p = ttest_rel(a_scores, b_scores) print(f"Paired t-test: t-value = {t_stat:.3f}, p-value = {t_p:.4f}") # Non-parametric alternative wilcox_stat, wilcox_p = wilcoxon(a_scores, b_scores, alternative='greater') print(f"Wilcoxon Signed-Rank Test (A > B): W-value = {wilcox_stat:.3f}, p-value = {wilcox_p:.4f}")
If the p-value is below your significance threshold (usually 0.05), you can conclude there's a significant difference. Compare the mean scores of A and B to confirm if A is better.
Scenario 2: Independent Samples (Different Students Took A vs B)
If Type A and Type B scores come from separate student groups (e.g., some years only used A, others used B, or different cohorts took different question types):
- First choice: Independent Samples t-test
Requires normality for both groups and equal variance (use Levene's test to check variance equality; if unequal, use Welch's t-test instead). - If normality fails: Mann-Whitney U Test
Non-parametric alternative for comparing independent groups.
Code Example
from scipy.stats import ttest_ind, mannwhitneyu, levene # Check variance equality levene_stat, levene_p = levene(a_scores, b_scores) equal_variance = levene_p > 0.05 # Independent t-test or Welch's test t_stat, t_p = ttest_ind(a_scores, b_scores, equal_var=equal_variance) print(f"Independent t-test: t-value = {t_stat:.3f}, p-value = {t_p:.4f}") # Non-parametric alternative (testing if A scores are higher) u_stat, u_p = mannwhitneyu(a_scores, b_scores, alternative='greater') print(f"Mann-Whitney U Test (A > B): U-value = {u_stat:.3f}, p-value = {u_p:.4f}")
Bonus: Account for Year-to-Year Variability
You have 10 years of data—don't ignore that time might affect scores (e.g., changing exam difficulty, curriculum shifts). To make your results more robust:
- Use Repeated Measures ANOVA if you have paired data across years, to control for year effects.
- Or use a Linear Mixed Model to account for both fixed effects (question type, year) and random effects (student/cohort variability).
Linear Mixed Model Example (Python with Statsmodels)
import pandas as pd import statsmodels.api as sm from statsmodels.formula.api import mixedlm # Convert your data to long format (assuming df has columns: student_id, year, a_score, b_score) df_long = df.melt( id_vars=['student_id', 'year'], value_vars=['a_score', 'b_score'], var_name='question_type', value_name='score' ) # Build model: question type and year as fixed effects, student_id as random effect model = mixedlm( formula="score ~ question_type + year", data=df_long, groups=df_long["student_id"] ) result = model.fit() print(result.summary())
Look for the question_type coefficient—if it's positive and statistically significant, that confirms Type A scores are higher while controlling for year and student differences.
Key Interpretation Tips
- Don't fixate only on p-values: calculate effect sizes (like Cohen's d for t-tests) to measure how large the difference between A and B really is (0.2 = small, 0.5 = medium, 0.8 = large).
- Visualize your data: use boxplots to compare score distributions, or line charts to show yearly average scores for both question types—this makes your findings easier to understand.
内容的提问来源于stack exchange,提问作者SummerRed

