You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

添加Positional Encoding后语言建模模型收敛性能恶化求助

位置编码导致语言模型性能骤降的问题排查请求

我正在基于两篇论文实现语言建模架构,参考了第一篇中的位置(时间)编码部分,同时结合了第二篇的相关内容。

以下是我的Keras核心实现代码:

word_seq = Input(shape = (SEQ_LEN,), dtype = "int32", name = "word_seq")
query = Input(shape = (EMBED_DIM, ), dtype = "float32", name = "q_input")
#the query for lang. modeling is a constant vector filled with 0.1, as described at the bottom of page 7 in the first paper
T_A = Added_Weights(input_dim = (SEQ_LEN, EMBED_DIM)) #Added_Weights is a custom layer I wrote, which I'll post below
#These are the "positional encoding" components
T_C = Added_Weights(input_dim = (SEQ_LEN, EMBED_DIM))
Emb_A = Embedding(output_dim = EMBED_DIM, input_dim = VOCAB_SIZE, input_length = SEQ_LEN, name = "Emb_A")
Emb_C = Embedding(output_dim = EMBED_DIM, input_dim = VOCAB_SIZE, input_length = SEQ_LEN, name = "Emb_C")
int_state_weights = Dense(units = EMBED_DIM, activation = 'linear', kernel_initializer=RandomNormal(mean=0., stddev = 0.05, seed = None))
layer_output = query #the loop uses the output from the previous layer as the query, but the first layer's query is just that constant vector

for i in range(0, NUM_LAYERS - 1):
    memories = Emb_A(word_seq) #these all re-use the weights instantiated earlier.
    memories = T_A(memories)
    memories = Dropout(DROPOUT_R)(memories)
    content = Emb_C(word_seq)
    content = T_C(content)
    mem_relevance = Dot(axes=[1, 2])([layer_output, memories])
    weighted_internal_state = int_state_weights(mem_relevance)
    mem_relevance = Softmax()(mem_relevance)
    content_relevance = Dot(axes=1)([mem_relevance, content]) # weight each piece of content by it's probability of being relevant
    layer_output = Add()([content_relevance, weighted_internal_state])
    layer_output = Dropout(DROPOUT_R)(layer_output)

final_output = Dense(units = VOCAB_SIZE, activation ='relu', kernel_initializer=RandomNormal(mean=0., stddev = 0.05, seed = None))(layer_output)
model = Model(inputs = [word_seq, query], outputs = prediction)
model.compile(optimizer = SGD(lr = 0.01, clipnorm = 50.), loss = 'categorical_crossentropy', metrics = ['accuracy'])
model.fit(x = [td_seqs, td_query], y = [td_labels], batch_size = BATCH_SIZE, callbacks = [lr_adjust, lr_termination, for_csv], epochs=200, verbose = 1)

当前BATCH_SIZE设为128,在未添加T_A和T_C(位置编码组件)时,使用约35000条训练样本训练后准确率可达96%。但添加位置编码后,训练结束时准确率仅约10%,训练损失约5.2。我将训练数据量扩大10倍后仍无明显改善。

我的Added_Weights自定义层代码如下:

class Added_Weights(Layer):
    def __init__(self, input_dim, **kwargs):
        super(Added_Weights, self).__init__(**kwargs)
        self.input_dim = input_dim

    def build(self, input_shape):
        # Create a trainable weight variable for this layer.
        self.kernel = self.add_weight(name='kernel',
                                      shape=(self.input_dim[0], self.input_dim[1]),
                                      initializer=RandomNormal(mean=0., stddev=0.05, seed=None),
                                      trainable=True)
        super(Added_Weights, self).build(input_shape)

    def call(self, x, **kwargs):
        return x + self.kernel

    def compute_output_shape(self, input_shape):
        return input_shape

查阅相关论文后明确位置编码应有效,但实际添加后模型性能大幅下降,恳请帮忙排查原因。

内容的提问来源于stack exchange,提问作者Sean Paulsen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:21:29