You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Vision Transformer模型绘制注意力图?项目实操求助

绘制Vision Transformer注意力图的实现方法

我在学校项目中实现Vision Transformer(ViT)模型,需要绘制注意力图以对比CNN模型与ViT模型的差异,但不确定具体操作方法。我使用的ViT模型为google/vit-base-patch16-224-in21k。

模型摘要

Model: "model"
_________________________________________________________________
 Layer (type)                Output Shape              Param #   
=================================================================
 input_1 (InputLayer)        [(None, 224, 224, 3)]     0         
                                                                 
 sequential (Sequential)     (None, 3, 224, 224)       0         
                                                                 
 vit (TFViTMainLayer)        TFBaseModelOutputWithPo   29686272  
                             oling(last_hidden_state              
                             =(None, 197, 768),                  
                              pooler_output=(None, 7              
                             68),                                
                              hidden_states=None, at              
                             tentions=None)                       
                                                                 
 tf.__operators__.getitem (  (None, 768)               0         
 SlicingOpLambda)                                                 
                                                                 
 dense (Dense)               (None, 2)                 1538      
                                                                 
=================================================================
Total params: 29687810 (113.25 MB)
Trainable params: 29687810 (113.25 MB)
Non-trainable params: 0 (0.00 Byte)
_________________________________________________________________

模型配置

ViTConfig {
  "_name_or_path": "google/vit-base-patch16-224-in21k",
  "attention_probs_dropout_prob": 0.0,
  "encoder_stride": 16,
  "hidden_act": "gelu",
  "hidden_dropout_prob": 0.0,
  "hidden_size": 768,
  "image_size": 224,
  "initializer_range": 0.02,
  "intermediate_size": 3072,
  "layer_norm_eps": 1e-12,
  "model_type": "vit",
  "num_attention_heads": 8,
  "num_channels": 3,
  "num_hidden_layers": 4,
  "patch_size": 16,
  "qkv_bias": true,
  "transformers_version": "4.38.2"
}

尝试过的代码

我尝试从原模型中提取激活层和输出层,但不确定如何重塑NumPy数组以让权重匹配224x224图像,当前代码没有得到需要的结果:

activation_layer = vit_model.get_layer("vit")
new_model = Model(inputs = vit_model.input, outputs = activation_layer.output)
final_dense = vit_model.get_layer('dense')
W = final_dense.get_weights()[0]

内容的提问来源于stack exchange,提问作者pokedot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 23:48:12