如何在将Keras层输出传入下一层前按指定方式打乱张量
解决Keras中逐样本独立打乱子序列的问题
问题本质
你要实现的是每个样本内部的子序列维度(第1轴)独立随机打乱,但tf.random.shuffle()默认只会打乱整个张量的第0轴(样本轴),直接使用必然不符合需求,甚至报错。
解决方案:自定义Lambda层实现逐样本打乱
下面是可直接复用的代码方案:
1. 定义逐样本打乱函数
这个函数会为每个样本生成独立的随机索引,再用索引重新排列子序列:
import tensorflow as tf from tensorflow.keras.layers import Lambda def shuffle_per_sample_subsequences(x): batch_size = tf.shape(x)[0] seq_len = tf.shape(x)[1] # 为每个样本生成0~seq_len-1的随机排列索引 shuffled_indices = tf.map_fn( lambda _: tf.random.shuffle(tf.range(seq_len)), tf.range(batch_size), dtype=tf.int32 ) # 按索引重排子序列,batch_dims=1保证每个样本用自己的索引打乱 return tf.gather(x, shuffled_indices, batch_dims=1)
2. 嵌入到你的模型中
在subtree_vectors生成后,插入这个Lambda层,替换后续Attention层的输入:
def create_model(embedding_weights, node_vocab_size, path_vocab_size, MAX_SUBTREE_LENGTH): config = Config() node_input = Input((MAX_SUBTREE_LENGTH,MAX_SUBTREE_LENGTH), dtype=tf.int32) path_input = Input((MAX_SUBTREE_LENGTH,), dtype=tf.int32) #embedding layer nodes_embedded = Embedding(node_vocab_size+2, config.embedding_size, trainable = True, name='node_embedding')(node_input) path_embedded = Embedding(path_vocab_size+2, config.embedding_size, trainable = True, name='path_embedding')(path_input) #(b,max_subtree,embedsize) # path embeddings from node embeddings nodes_embedded_merged = K.sum(nodes_embedded, axis=2) #(b,max_subtree,embedsize) node_path_merged = concatenate([nodes_embedded_merged, path_embedded]) subtree_vectors = TimeDistributed(Dense(config.embedding_size*2, use_bias=False, activation='tanh'))(node_path_merged) # ------------------- 插入打乱逻辑 ------------------- subtree_vectors = Lambda(shuffle_per_sample_subsequences)(subtree_vectors) # -------------------------------------------------- # Attention Layer attention_vectors = Dense(1,)(subtree_vectors) attention_weights = Softmax(axis=1)(attention_vectors) # Generating code vectors code_vectors = K.sum(subtree_vectors * attention_weights, axis=1) # Prediction layer output_class = Dense(config.num_classes, use_bias=False, activation='softmax')(code_vectors) model = Model(inputs=[node_input, path_input], outputs=output_class) return model
额外说明
- 训练时每次前向传播都会生成不同的打乱顺序,满足随机增强需求;
- 如果需要在测试阶段固定顺序(不打乱),可以给函数加
training参数控制:
def shuffle_per_sample_subsequences(x, training=True): if not training: return x batch_size = tf.shape(x)[0] seq_len = tf.shape(x)[1] shuffled_indices = tf.map_fn( lambda _: tf.random.shuffle(tf.range(seq_len)), tf.range(batch_size), dtype=tf.int32 ) return tf.gather(x, shuffled_indices, batch_dims=1)
然后在Lambda层传入参数:
subtree_vectors = Lambda(shuffle_per_sample_subsequences, arguments={'training': True})(subtree_vectors)
训练时传True,测试时改False即可。
内容的提问来源于stack exchange,提问作者mhoq
相关产品推荐
相关产品推荐

