为何SAS实现的Cochran-Mantel-Haenszel检验与其他版本存在差异?
Great question—this is a super common gotcha when comparing CMH results across statistical tools, especially when mixing ordered and unordered factors like your quantile groups (Q1-Q4) and categorical groups (A/B/C). Let’s break down the key reasons for the discrepancies:
1. SAS Automatically Adapts to Ordered Variables (Others Often Don’t)
SAS’s PROC FREQ with the CMH option doesn’t just run a single CMH test—it calculates three distinct statistics tailored to different variable types:
- General Association: The standard CMH chi-square, designed for two unordered factors (matches the default output of tools like R’s
mantelhaen.test()or SPSS’s basic CMH). - Row Mean Scores Differ: A specialized statistic for when one factor is ordered (your Q1-Q4) and the other is unordered (your groups A/B/C). SAS prioritizes this result when it detects an ordered variable, since it’s more powerful for testing trends across quantiles.
- Nonzero Correlation: For when both factors are ordered (not your case, but included for completeness).
Most other tools default to only running the General Association statistic unless you explicitly tell them to account for ordered variables. That’s the biggest source of mismatch—you’re likely comparing SAS’s Row Mean Scores result to another tool’s default General Association result.
2. Different Default Scoring for Ordered Factors
When calculating the ordered-specific CMH statistic, SAS uses rank scores for your quantile variable by default. Other tools might use simpler integer scores (1 for Q1, 2 for Q2, etc.) or require you to manually specify the scoring method. Even small differences in how ordered categories are weighted can shift the test statistic and p-value.
3. How Tools Handle Input Data
If you’re feeding raw data to SAS vs. precomputed contingency tables to another tool, double-check how each tool aggregates counts. SAS’s PROC FREQ automatically handles duplicate rows as frequencies, but some tools might require you to explicitly weight rows by count, leading to errors if you miss this step.
Example: Aligning Results Across Tools
To get matching results:
- If you want SAS to match another tool’s default CMH: Look at the General Association statistic in SAS’s
PROC FREQoutput, not the Row Mean Scores one. - If you want another tool to match SAS’s ordered CMH:
- In R, use the
epitools::cmh.test()function and specifyscore = "ranks"to mirror SAS’s scoring:library(epitools) cmh_result <- cmh.test(table(your_data$group, your_data$quantile), score = "ranks") print(cmh_result$statistic["row.mean.score"]) - This will give you the same statistic as SAS’s Row Mean Scores Differ output.
- In R, use the
Key Takeaway
The core issue is that SAS is proactive about leveraging the ordered nature of your quantile variable, while most other tools stick to a one-size-fits-all CMH test by default. Always check which specific CMH statistic each tool is calculating, and adjust settings to match the variable types you’re working with.
内容的提问来源于stack exchange,提问作者Dan Chaltiel

