You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于GLMM利用成绩数据判定不同题型的表现差异

Hey there! Let's walk through how to figure out if students perform better on Type A questions compared to Type B, using your 10 years of exam data. I'll break down the best analysis approaches based on your data structure, plus some extra considerations for that long-term dataset.

Core Analysis Strategies

Your goal is to test if the average score on Type A is significantly higher than Type B. Since scores are continuous (0-100), we'll start with parametric tests (when assumptions hold) and use non-parametric alternatives if needed.

Scenario 1: Paired Data (Same Students Took Both Question Types)

If every student completed both Type A and Type B questions in each exam, your data is paired (two scores per student). This is the most common case for comparing two question types in the same test:

  • First choice: Paired t-test
    Requires that the difference between each student's A and B scores follows a normal distribution. Verify this with a Shapiro-Wilk test or QQ plot.
  • If normality fails: Wilcoxon Signed-Rank Test
    A non-parametric alternative that doesn't rely on normal distribution assumptions.

Code Example (Python with Scipy)

from scipy.stats import ttest_rel, wilcoxon, shapiro

# Assume a_scores and b_scores are arrays of corresponding student scores
# First check normality of differences
diff_scores = [a - b for a, b in zip(a_scores, b_scores)]
shapiro_stat, shapiro_p = shapiro(diff_scores)
print(f"Shapiro-Wilk normality test p-value: {shapiro_p:.4f}")

# Paired t-test
t_stat, t_p = ttest_rel(a_scores, b_scores)
print(f"Paired t-test: t-value = {t_stat:.3f}, p-value = {t_p:.4f}")

# Non-parametric alternative
wilcox_stat, wilcox_p = wilcoxon(a_scores, b_scores, alternative='greater')
print(f"Wilcoxon Signed-Rank Test (A > B): W-value = {wilcox_stat:.3f}, p-value = {wilcox_p:.4f}")

If the p-value is below your significance threshold (usually 0.05), you can conclude there's a significant difference. Compare the mean scores of A and B to confirm if A is better.

Scenario 2: Independent Samples (Different Students Took A vs B)

If Type A and Type B scores come from separate student groups (e.g., some years only used A, others used B, or different cohorts took different question types):

  • First choice: Independent Samples t-test
    Requires normality for both groups and equal variance (use Levene's test to check variance equality; if unequal, use Welch's t-test instead).
  • If normality fails: Mann-Whitney U Test
    Non-parametric alternative for comparing independent groups.

Code Example

from scipy.stats import ttest_ind, mannwhitneyu, levene

# Check variance equality
levene_stat, levene_p = levene(a_scores, b_scores)
equal_variance = levene_p > 0.05

# Independent t-test or Welch's test
t_stat, t_p = ttest_ind(a_scores, b_scores, equal_var=equal_variance)
print(f"Independent t-test: t-value = {t_stat:.3f}, p-value = {t_p:.4f}")

# Non-parametric alternative (testing if A scores are higher)
u_stat, u_p = mannwhitneyu(a_scores, b_scores, alternative='greater')
print(f"Mann-Whitney U Test (A > B): U-value = {u_stat:.3f}, p-value = {u_p:.4f}")

Bonus: Account for Year-to-Year Variability

You have 10 years of data—don't ignore that time might affect scores (e.g., changing exam difficulty, curriculum shifts). To make your results more robust:

  • Use Repeated Measures ANOVA if you have paired data across years, to control for year effects.
  • Or use a Linear Mixed Model to account for both fixed effects (question type, year) and random effects (student/cohort variability).

Linear Mixed Model Example (Python with Statsmodels)

import pandas as pd
import statsmodels.api as sm
from statsmodels.formula.api import mixedlm

# Convert your data to long format (assuming df has columns: student_id, year, a_score, b_score)
df_long = df.melt(
    id_vars=['student_id', 'year'],
    value_vars=['a_score', 'b_score'],
    var_name='question_type',
    value_name='score'
)

# Build model: question type and year as fixed effects, student_id as random effect
model = mixedlm(
    formula="score ~ question_type + year",
    data=df_long,
    groups=df_long["student_id"]
)
result = model.fit()
print(result.summary())

Look for the question_type coefficient—if it's positive and statistically significant, that confirms Type A scores are higher while controlling for year and student differences.

Key Interpretation Tips

  • Don't fixate only on p-values: calculate effect sizes (like Cohen's d for t-tests) to measure how large the difference between A and B really is (0.2 = small, 0.5 = medium, 0.8 = large).
  • Visualize your data: use boxplots to compare score distributions, or line charts to show yearly average scores for both question types—this makes your findings easier to understand.

内容的提问来源于stack exchange,提问作者SummerRed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:01:51