添加Positional Encoding后语言建模模型收敛性能恶化求助
位置编码导致语言模型性能骤降的问题排查请求
我正在基于两篇论文实现语言建模架构,参考了第一篇中的位置(时间)编码部分,同时结合了第二篇的相关内容。
以下是我的Keras核心实现代码:
word_seq = Input(shape = (SEQ_LEN,), dtype = "int32", name = "word_seq") query = Input(shape = (EMBED_DIM, ), dtype = "float32", name = "q_input") #the query for lang. modeling is a constant vector filled with 0.1, as described at the bottom of page 7 in the first paper T_A = Added_Weights(input_dim = (SEQ_LEN, EMBED_DIM)) #Added_Weights is a custom layer I wrote, which I'll post below #These are the "positional encoding" components T_C = Added_Weights(input_dim = (SEQ_LEN, EMBED_DIM)) Emb_A = Embedding(output_dim = EMBED_DIM, input_dim = VOCAB_SIZE, input_length = SEQ_LEN, name = "Emb_A") Emb_C = Embedding(output_dim = EMBED_DIM, input_dim = VOCAB_SIZE, input_length = SEQ_LEN, name = "Emb_C") int_state_weights = Dense(units = EMBED_DIM, activation = 'linear', kernel_initializer=RandomNormal(mean=0., stddev = 0.05, seed = None)) layer_output = query #the loop uses the output from the previous layer as the query, but the first layer's query is just that constant vector for i in range(0, NUM_LAYERS - 1): memories = Emb_A(word_seq) #these all re-use the weights instantiated earlier. memories = T_A(memories) memories = Dropout(DROPOUT_R)(memories) content = Emb_C(word_seq) content = T_C(content) mem_relevance = Dot(axes=[1, 2])([layer_output, memories]) weighted_internal_state = int_state_weights(mem_relevance) mem_relevance = Softmax()(mem_relevance) content_relevance = Dot(axes=1)([mem_relevance, content]) # weight each piece of content by it's probability of being relevant layer_output = Add()([content_relevance, weighted_internal_state]) layer_output = Dropout(DROPOUT_R)(layer_output) final_output = Dense(units = VOCAB_SIZE, activation ='relu', kernel_initializer=RandomNormal(mean=0., stddev = 0.05, seed = None))(layer_output) model = Model(inputs = [word_seq, query], outputs = prediction) model.compile(optimizer = SGD(lr = 0.01, clipnorm = 50.), loss = 'categorical_crossentropy', metrics = ['accuracy']) model.fit(x = [td_seqs, td_query], y = [td_labels], batch_size = BATCH_SIZE, callbacks = [lr_adjust, lr_termination, for_csv], epochs=200, verbose = 1)
当前BATCH_SIZE设为128,在未添加T_A和T_C(位置编码组件)时,使用约35000条训练样本训练后准确率可达96%。但添加位置编码后,训练结束时准确率仅约10%,训练损失约5.2。我将训练数据量扩大10倍后仍无明显改善。
我的Added_Weights自定义层代码如下:
class Added_Weights(Layer): def __init__(self, input_dim, **kwargs): super(Added_Weights, self).__init__(**kwargs) self.input_dim = input_dim def build(self, input_shape): # Create a trainable weight variable for this layer. self.kernel = self.add_weight(name='kernel', shape=(self.input_dim[0], self.input_dim[1]), initializer=RandomNormal(mean=0., stddev=0.05, seed=None), trainable=True) super(Added_Weights, self).build(input_shape) def call(self, x, **kwargs): return x + self.kernel def compute_output_shape(self, input_shape): return input_shape
查阅相关论文后明确位置编码应有效,但实际添加后模型性能大幅下降,恳请帮忙排查原因。
内容的提问来源于stack exchange,提问作者Sean Paulsen
相关产品推荐
相关产品推荐

