如何解析TraMineR包seqecmpgroup函数输出中的残差?
seqecmpgroup输出的残差(Residuals)含义与作用 Great question! Let's unpack what those residual columns (like Resid.Low, Resid.Medium, Resid.High) in the seqecmpgroup output mean and how they help you interpret group differences in sequence patterns.
一、残差的核心含义
First, remember that seqecmpgroup runs a chi-squared test to compare the frequency of subsequences across your groups (Low/Medium/High in your example). The residuals here are Pearson residuals, which standardize the difference between the observed frequency of a subsequence in a group and the expected frequency (what you'd see if there was no difference between groups).
Here's the breakdown:
- Calculation logic:
Residual = (Observed Frequency - Expected Frequency) / sqrt(Expected Frequency * (1 - Expected Frequency) * (Group Size / Total Sample Size)) - Sign interpretation:
- A positive residual means the subsequence appears more often in that group than expected if groups were identical.
- A negative residual means the subsequence appears less often in that group than expected.
- Magnitude: The larger the absolute value of the residual, the bigger the gap between observed and expected frequency. This aligns with the
p.valuecolumn—residuals with large absolute values usually correspond to small p-values (statistically significant differences).
Using your sample table for example:
For the subsequence
A-B,Resid.Lowis -2.98 andResid.Highis 2.61. This tells us that in the Low group,A-Boccurs way less often than expected (observed freq = 0.1619), while in the High group, it occurs way more often than expected (observed freq = 0.3619).
二、残差的实际作用
Residuals add critical context beyond just knowing a subsequence has a statistically significant group difference (from the p.value):
- Pinpoint the direction of group differences: The p-value tells you "this subsequence differs across groups," but residuals tell you which groups have more/less of it. For
A-C-D, you can see Low groups underperform and High groups overperform relative to expectations. - Prioritize meaningful patterns: When you have a long list of subsequences, focus on those with small p-values and large absolute residuals—these are the patterns that most strongly distinguish your groups.
- Interpret sequence behavior: Residuals help you tell a story about your groups. For instance,
C-D-Fhas a negative residual in Low groups and positive residuals in Medium/High, suggesting this subsequence becomes more common as you move from Low to High groups.
内容的提问来源于stack exchange,提问作者user3315563

