基于Luong注意力模型的英中翻译模型体积异常过大问题排查
Problem Description
I'm implementing an English-Chinese translation model based on Minh-Thang Luong's attention model. After training, the model size is unexpectedly large (980MB). According to Luong's original paper, the model parameters are:
- State size: 120
- Source language vocabulary size: 400,000
- Source language word embedding dimension: 400,000*50
- Target language vocabulary size: 20,000
- Target language word embedding dimension: 20,000*300
Here's my TensorFlow implementation code:
import tensorflow as tf src_vocab_size=400000 src_w2v_dim=50 tgt_vocab_size=20000 tgt_w2v_dim=300 state_size=120 with tf.variable_scope('net_encode'): ph_src_embedding = tf.placeholder(dtype=tf.float32,shape=[src_vocab_size,src_w2v_dim],name='src_vocab_embedding_placeholder') #src_word_emb = tf.Variable(initial_value=ph_src_embedding,dtype=tf.float32,trainable=False, name='src_vocab_embedding_variable') encoder_X_ix = tf.placeholder(shape=(None, None), dtype=tf.int32) encoder_X_len = tf.placeholder(shape=(None), dtype=tf.int32) encoder_timestep = tf.shape(encoder_X_ix)[1] encoder_X = tf.nn.embedding_lookup(ph_src_embedding, encoder_X_ix) batchsize = tf.shape(encoder_X_ix)[0] encoder_Y_ix = tf.placeholder(shape=[None, None],dtype=tf.int32) encoder_Y_onehot = tf.one_hot(encoder_Y_ix, src_vocab_size) enc_cell = tf.nn.rnn_cell.LSTMCell(state_size) enc_initstate = enc_cell.zero_state(batchsize,dtype=tf.float32) enc_outputs, enc_final_states = tf.nn.dynamic_rnn(enc_cell,encoder_X,encoder_X_len,enc_initstate) enc_pred = tf.layers.dense(enc_outputs, units=src_vocab_size) encoder_loss = tf.losses.softmax_cross_entropy(encoder_Y_onehot,enc_pred) encoder_trainop = tf.train.AdamOptimizer(0.001).minimize(encoder_loss) with tf.variable_scope('net_decode'): ph_tgt_embedding = tf.placeholder(dtype=tf.float32, shape=[tgt_vocab_size, tgt_w2v_dim], name='tgt_vocab_embedding_placeholder') #tgt_word_emb = tf.Variable(initial_value=ph_tgt_embedding, dtype=tf.float32, trainable=False, name='tgt_vocab_embedding_variable') decoder_X_ix = tf.placeholder(shape=(None, None), dtype=tf.int32) decoder_timestep = tf.shape(decoder_X_ix)[1] decoder_X_len = tf.placeholder(shape=(None), dtype=tf.int32) decoder_X = tf.nn.embedding_lookup(ph_tgt_embedding, decoder_X_ix) decoder_Y_ix = tf.placeholder(shape=[None, None],dtype=tf.int32) decoder_Y_onehot = tf.one_hot(decoder_Y_ix, tgt_vocab_size) dec_cell = tf.nn.rnn_cell.LSTMCell(state_size) dec_outputs, dec_final_state = tf.nn.dynamic_rnn(dec_cell,decoder_X,decoder_X_len,enc_final_states) tile_enc = tf.tile(tf.expand_dims(enc_outputs,1),[1,decoder_timestep,1,1]) # [batchsize,decoder_len,encoder_len,state_size] tile_dec = tf.tile(tf.expand_dims(dec_outputs, 2), [1, 1, encoder_timestep, 1]) # [batchsize,decoder_len,encoder_len,state_size] enc_dec_cat = tf.concat([tile_enc,tile_dec],-1) # [batchsize,decoder_len,encoder_len,state_size*2] weights = tf.nn.softmax(tf.layers.dense(enc_dec_cat,units=1),axis=-2) # [batchsize,decoder_len,encoder_len,1] weighted_enc = tf.tile(weights, [1, 1, 1, state_size])*tf.tile(tf.expand_dims(enc_outputs,1),[1,decoder_timestep,1,1]) # [batchsize,decoder_len,encoder_len,state_size] attention = tf.reduce_sum(weighted_enc,axis=2,keepdims=False) # [batchsize,decoder_len,state_size] dec_attention_cat = tf.concat([dec_outputs,attention],axis=-1) # [batchsize,decoder_len,state_size*2] dec_pred = tf.layers.dense(dec_attention_cat,units=tgt_vocab_size) # [batchsize,decoder_len,tgt_vocab_size] pred_ix = tf.argmax(dec_pred,axis=-1) # [batchsize,decoder_len] decoder_loss = tf.losses.softmax_cross_entropy(decoder_Y_onehot,dec_pred) total_loss = encoder_loss + decoder_loss decoder_trainop = tf.train.AdamOptimizer(0.001).minimize(total_loss) _l0 = tf.summary.scalar('decoder_loss',decoder_loss) _l1 = tf.summary.scalar('encoder_loss',encoder_loss) log_all = tf.summary.merge_all() writer = tf.summary.FileWriter(log_path,graph=tf.get_default_graph())
My manual calculation of model parameter size is as follows:
- Encoder cell:
(50*120+120*120+120)*4 = 82080floats - Encoder dense layer:
120*400000 = 48000000floats - Decoder cell:
(300*120+120*120+120)*4 = 202080floats - Dense layer for attention weights:
(120+120)*1 = 240floats - Decoder dense layer:
(120+120)*20000 = 4800000floats
Total converted size is around 212MB, but the actual model size reaches 980MB. Where's the problem?
Answer
Hey there, let's break down why your model is way larger than expected. The main culprits are optimizer state duplication from using two separate Adam optimizers and a few small oversights in your parameter count. Here's the detailed breakdown:
1. The Biggest Issue: Dual Adam Optimizers Create Duplicate State
Your code defines two separate tf.train.AdamOptimizer instances: one for encoder_trainop and another for decoder_trainop.
Adam optimizers track two extra values per trainable parameter (first-order momentum m and second-order momentum v) to adapt learning rates dynamically. When you use two optimizers:
- The first optimizer (for encoder loss) saves state only for encoder-related variables.
- The second optimizer (for total loss) saves state for all trainable variables (encoder + decoder).
This leads to massive state bloat:
- Encoder variables end up with 3x their original size (original parameters + 2 sets of optimizer state).
- Decoder variables end up with 3x their original size (original parameters + 1 set of optimizer state).
Using your ~212MB base parameter size as a starting point:
- Original params: ~212MB
- First optimizer state (encoder-only): ~387MB
- Second optimizer state (all params): ~424MB
- Total estimated size: ~1023MB, which aligns closely with your reported 980MB (the small difference comes from rounding and tiny extra variables like optimizer beta parameters).
2. Small Parameter Count Oversights
Your manual calculation missed bias terms in dense layers, which add a small but measurable amount to the size:
- Encoder dense layer (
enc_pred): You calculated120*400000, but dense layers include a bias term for each output unit, adding 400,000 extra floats (~1.6MB). - Decoder attention dense layer: You calculated
(120+120)*1, but missed the single bias term (negligible size, but still an oversight). - Decoder final dense layer (
dec_pred): Missed 20,000 bias floats (~80KB).
These are minor compared to the optimizer state issue, but they contribute to the discrepancy.
3. TensorFlow Checkpoint/SavedModel Overhead
If you're using tf.train.Saver or SavedModel, the saved directory includes extra files:
.meta: Stores the computation graph structure (usually a few MBs)..index: Maps variable names to data locations (tiny size)..data: The actual variable values (the bulk of the size).
This overhead is small compared to the optimizer state problem, but it adds a little extra size.
Fixes to Reduce Model Size
- Use a single optimizer: Combine your training operations into one optimizer that minimizes the total loss. This eliminates the duplicate optimizer state from the second Adam instance:
# Replace both train ops with this single block optimizer = tf.train.AdamOptimizer(0.001) total_loss = encoder_loss + decoder_loss train_op = optimizer.minimize(total_loss) - Freeze the model for inference: If you only need the model for translation (not further training), freeze it to strip out optimizer state and non-inference nodes. Use
tf.graph_util.convert_variables_to_constantsto create a single frozen graph file. - Verify variable scopes: Ensure you're not accidentally creating duplicate variables (your code uses separate
net_encodeandnet_decodescopes, so this isn't an issue here).
Implementing these fixes should bring your model size down to something much closer to your original ~212MB calculation!
内容的提问来源于stack exchange,提问作者abracadabra

