Kaldi入门教程:如何使用tri2a/tri2b_mmi训练并替换tri1方法
Hey there! Let's tackle your Kaldi questions one by one—super glad you're experimenting with custom data and advanced training setups.
The Kaldi for Dummies tutorial sticks to the foundational mono + tri1 pipeline, but extending it to tri2a (LDA+MLLT trained model) and tri2b_mmi (MMI discriminative model) is totally doable. Here's how to slot these steps into your existing workflow:
Step 1: Train the tri2a (LDA+MLLT) model
Assuming you've already completed the mono alignment and tri1 training steps from the tutorial, run the LDA+MLLT training script. This optimizes the feature space and refines the triphone model:
steps/train_lda_mllt.sh --cmd "$train_cmd" 2500 15000 data/train data/lang exp/tri1 exp/tri2a
2500: Number of base GMM components (adjust based on your dataset size)15000: Total number of mixture componentsexp/tri1: Path to your existing tri1 model directoryexp/tri2a: Output directory for the tri2a model
Step 2: Align data with tri2a
You need aligned data to train the discriminative tri2b_mmi model. Use the tri2a model to re-align your training set:
steps/align_si.sh --cmd "$train_cmd" --use-graphs true data/train data/lang exp/tri2a exp/tri2a_ali
exp/tri2a_ali: Directory to store tri2a-aligned training data
Step 3: Train tri2b_mmi (MMI discriminative model)
Now run the MMI training, which fine-tunes the model by maximizing mutual information between the acoustic model and true transcriptions:
steps/train_mmi.sh --boost 0.1 data/train data/lang exp/tri2a_ali exp/tri2a exp/tri2b_mmi
--boost 0.1: Boost factor for correct phones (helps stabilize discriminative training)exp/tri2b_mmi: Output directory for the final MMI model
Step 4: Decode with tri2b_mmi
Update the tutorial's decoding step to use your new tri2b_mmi model for better accuracy:
steps/decode.sh --cmd "$decode_cmd" exp/tri2b_mmi/graph data/test exp/tri2b_mmi/decode_test
Short answer: You can't fully skip the mono and tri1 steps directly, but you can extend your existing pipeline to use tri2a/tri2b_mmi instead of stopping at tri1. Here's why:
- The tri2a model relies on the tri1 model's triphone alignments and initial GMM parameters to compute the LDA+MLLT feature transformations. Without tri1, there's no foundational triphone model to build from.
- Mono alignment is the first step to get rough phone alignments for training the initial triphone model (tri1)—it's the mandatory starting point for all subsequent triphone-based training in Kaldi.
That said, if you've already completed the mono and tri1 steps with your custom data, you can absolutely replace the tutorial's post-tri1 steps with the tri2a + tri2b_mmi workflow outlined above. This will give you a more accurate, discriminative model compared to just stopping at tri1.
内容的提问来源于stack exchange,提问作者Geo George

