关于Forward-Backward算法起始与终止状态转移概率估计的疑问
Great question—this is a super common point of confusion when diving into HMM training, since different experts frame start/end state handling in subtly different ways. Let’s break down why Jurafsky & Martin, Jason Eisner, and Rabiner’s tutorials seem to disagree:
1. Jurafsky & Martin’s Core Framework
Jurafsky and Martin’s textbook frames the HMM with the classic three-component model: initial state distribution π, transition matrix A (for state-to-state transitions between observation steps), and observation likelihood matrix B.
When they say the Forward-Backward algorithm doesn’t estimate start/end transition probabilities, they mean:
- The "start" is modeled by
π(the probability of being in each state at time step 1), not as a transition from a virtual start state. - The "end" is treated as the natural conclusion of the sequence—there’s no explicit transition to a virtual end state; instead, we use the final forward probabilities to compute the sequence likelihood.
In this setup, A only covers transitions between actual observation states, so start/end-related probabilities live outside the transition matrix and aren’t estimated via the same Forward-Backward steps used for A.
2. Jason Eisner’s Extended Virtual State Model
Eisner’s tables explicitly introduce virtual start (S) and end (E) states to model the entire sequence as a complete path: S → state_1 → state_2 → ... → state_T → E.
In this framework:
- The transition from
Sto the first real state becomes part of the transition matrixA. - The transition from the last real state to
Eis also part ofA.
Since these are now formal entries in the transition matrix, the Forward-Backward algorithm does estimate these probabilities alongside all other state transitions. This makes the sequence’s start and end feel like integral parts of the transition process, which can be helpful for teaching or certain types of sequence modeling tasks.
3. Rabiner’s Classic Tutorial Approach
Rabiner’s iconic HMM tutorial takes a middle ground. He focuses on the initial state distribution π (the probability of starting in each state) as the "start" component, similar to Jurafsky & Martin. However, he sometimes refers to π informally as the "probability of transitioning from the start state"—but this doesn’t mean he’s using a virtual start state.
Rabiner’s framework doesn’t introduce a virtual end state at all, so there’s no need to estimate transitions to an end state. The sequence ends when we reach the final observation step, and we use the forward probabilities at that step to compute the overall likelihood.
Key Takeaway
All three approaches are consistent at their core—they just use different modeling conventions for start/end behavior:
- If you use virtual start/end states (Eisner), those transitions become part of
Aand get estimated by Forward-Backward. - If you use an initial distribution
πand no virtual end state (Jurafsky & Martin, Rabiner), those start/end probabilities are separate parameters, not part of the transition matrixAthat Forward-Backward estimates.
内容的提问来源于stack exchange,提问作者Sela Fried

