如何对注意力操作进行向量化以避免Python for循环提升训练速度
注意力循环的向量化改造方案
完全可以通过广播机制调整张量维度,彻底移除Python for循环,大幅提升运算效率。你提到的直接传入高维张量到AttentionNet是可行的:nn.Sequential封装的网络默认仅对最后一维做运算,前面所有维度都会被识别为批量维度,只要保证最后一维长度是2*E即可。
具体改造步骤
- 首先对两个嵌入张量做维度扩展,利用广播机制避免重复拷贝:
原embeddings形状为(B, E),扩展为(B, I, E);原other_embeddings形状为(I, E),扩展为(B, I, E) - 沿最后一维拼接两个扩展后的张量,得到所有配对的嵌入组合,形状为
(B, I, 2*E) - 直接将拼接后的张量传入
AttentionNet,得到形状为(B, I, 1)的得分,压缩最后一维即可得到和原来循环完全一致的(B, I)注意力得分
改造后代码示例
def forward(self, input_features, ...): ... """ B = batch size, I = number of items for attention, E = embedding size """ ... # get embeddings from input features for current batch embeddings = self.embedding_layer(input_features) # (B, E) other_embeddings = self.embedding_layer(other_features) # (I, E) # 向量化计算注意力得分,完全移除Python循环 # 维度扩展 emb_expanded = embeddings.detach().unsqueeze(1).expand(-1, I, -1) # (B, I, E) other_emb_expanded = other_embeddings.detach().unsqueeze(0).expand(B, -1, -1) # (B, I, E) # 拼接所有配对 item_emb_pairs = torch.cat([emb_expanded, other_emb_expanded], dim=-1) # (B, I, 2*E) # 批量计算所有得分 attention_scores = self.AttentionNet(item_emb_pairs).squeeze(-1) # (B, I) # pass through softmax attention_scores = F.softmax(attention_scores, dim=1) # (B, I) # 后续加权求和也可以直接向量化完成,不需要循环 output = torch.matmul(attention_scores.unsqueeze(1), other_emb_expanded).squeeze(1) # (B, E) ...
补充说明
如果你习惯使用(I, B, 2*E)的维度布局也完全可行,只需调整扩展时的维度顺序,计算完成后将得分转置为(B, I)即可,不会影响计算效率。
内容的提问来源于stack exchange,提问作者Michael
相关产品推荐
相关产品推荐

