You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Forward-Backward算法起始与终止状态转移概率估计的疑问

Understanding the Forward-Backward Algorithm's Start/End State Discrepancy

Great question—this is a super common point of confusion when diving into HMM training, since different experts frame start/end state handling in subtly different ways. Let’s break down why Jurafsky & Martin, Jason Eisner, and Rabiner’s tutorials seem to disagree:

1. Jurafsky & Martin’s Core Framework

Jurafsky and Martin’s textbook frames the HMM with the classic three-component model: initial state distribution π, transition matrix A (for state-to-state transitions between observation steps), and observation likelihood matrix B.

When they say the Forward-Backward algorithm doesn’t estimate start/end transition probabilities, they mean:

  • The "start" is modeled by π (the probability of being in each state at time step 1), not as a transition from a virtual start state.
  • The "end" is treated as the natural conclusion of the sequence—there’s no explicit transition to a virtual end state; instead, we use the final forward probabilities to compute the sequence likelihood.

In this setup, A only covers transitions between actual observation states, so start/end-related probabilities live outside the transition matrix and aren’t estimated via the same Forward-Backward steps used for A.

2. Jason Eisner’s Extended Virtual State Model

Eisner’s tables explicitly introduce virtual start (S) and end (E) states to model the entire sequence as a complete path: S → state_1 → state_2 → ... → state_T → E.

In this framework:

  • The transition from S to the first real state becomes part of the transition matrix A.
  • The transition from the last real state to E is also part of A.

Since these are now formal entries in the transition matrix, the Forward-Backward algorithm does estimate these probabilities alongside all other state transitions. This makes the sequence’s start and end feel like integral parts of the transition process, which can be helpful for teaching or certain types of sequence modeling tasks.

3. Rabiner’s Classic Tutorial Approach

Rabiner’s iconic HMM tutorial takes a middle ground. He focuses on the initial state distribution π (the probability of starting in each state) as the "start" component, similar to Jurafsky & Martin. However, he sometimes refers to π informally as the "probability of transitioning from the start state"—but this doesn’t mean he’s using a virtual start state.

Rabiner’s framework doesn’t introduce a virtual end state at all, so there’s no need to estimate transitions to an end state. The sequence ends when we reach the final observation step, and we use the forward probabilities at that step to compute the overall likelihood.

Key Takeaway

All three approaches are consistent at their core—they just use different modeling conventions for start/end behavior:

  • If you use virtual start/end states (Eisner), those transitions become part of A and get estimated by Forward-Backward.
  • If you use an initial distribution π and no virtual end state (Jurafsky & Martin, Rabiner), those start/end probabilities are separate parameters, not part of the transition matrix A that Forward-Backward estimates.

内容的提问来源于stack exchange,提问作者Sela Fried

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:17:18