You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化PyTorch Transformer模型以实现CPU端更快推理?

CPU端Transformer模型推理速度优化求助

我希望大幅提升某Transformer模型在CPU端的推理速度,当前该模型单次推理耗时4.5秒,我使用torchinfo库中的summary(self.model, input_size=torch.Size([1, 256, 192]))方法完成了模型结构分析。

模型代码出自sony的hFT-Transformer项目中的model_spec2midi.py。

模型结构分析结果

=========================================================================================================
Layer (type:depth-idx)                                  Output Shape              Param #
=========================================================================================================
Model_SPEC2MIDI                                         [1, 128, 88]              --
├─Encoder_SPEC2MIDI: 1-1                                [1, 128, 256, 256]        --
│    └─Conv2d: 2-1                                      [128, 4, 256, 61]         24
│    └─Linear: 2-2                                      [128, 256, 256]           62,720
│    └─Embedding: 2-3                                   [128, 256, 256]           65,536
│    └─Dropout: 2-4                                     [128, 256, 256]           --
│    └─ModuleList: 2-5                                  --                        --
│    │    └─EncoderLayer: 3-1                           [128, 256, 256]           526,592
│    │    └─EncoderLayer: 3-2                           [128, 256, 256]           526,592
│    │    └─EncoderLayer: 3-3                           [128, 256, 256]           526,592
├─Decoder_SPEC2MIDI: 1-2                                [1, 128, 88]              --
│    └─Embedding: 2-6                                   [128, 88, 256]            22,528
│    └─DecoderLayer_Zero: 2-7                           [128, 88, 256]            --
│    │    └─MultiHeadAttentionLayer: 3-4                [128, 88, 256]            263,168
│    │    └─Dropout: 3-5                                [128, 88, 256]            --
│    │    └─LayerNorm: 3-6                              [128, 88, 256]            512
│    │    └─PositionwiseFeedforwardLayer: 3-7           [128, 88, 256]            262,912
│    │    └─Dropout: 3-8                                [128, 88, 256]            --
│    │    └─LayerNorm: 3-9                              [128, 88, 256]            (recursive)
│    └─ModuleList: 2-8                                  --                        --
│    │    └─DecoderLayer: 3-10                          [128, 88, 256]            789,760
│    │    └─DecoderLayer: 3-11                          [128, 88, 256]            789,760
│    └─Linear: 2-9                                      [128, 88, 1]              257
│    └─Sigmoid: 2-10                                    [1, 128, 88]              --
│    └─Linear: 2-11                                     [128, 88, 1]              257
│    └─Sigmoid: 2-12                                    [1, 128, 88]              --
│    └─Linear: 2-13                                     [128, 88, 1]              257
│    └─Sigmoid: 2-14                                    [1, 128, 88]              --
│    └─Linear: 2-15                                     [128, 88, 128]            32,896
│    └─Embedding: 2-16                                  [88, 128, 256]            32,768
│    └─Dropout: 2-17                                    [88, 128, 256]            --
│    └─ModuleList: 2-18                                 --                        --
│    │    └─EncoderLayer: 3-12                          [88, 128, 256]            526,592
│    │    └─EncoderLayer: 3-13                          [88, 128, 256]            526,592
│    │    └─EncoderLayer: 3-14                          [88, 128, 256]            526,592
│    └─Linear: 2-19                                     [88, 128, 1]              257
│    └─Sigmoid: 2-20                                    [1, 128, 88]              --
│    └─Linear: 2-21                                     [88, 128, 1]              257
│    └─Sigmoid: 2-22                                    [1, 128, 88]              --
│    └─Linear: 2-23                                     [88, 128, 1]              257
│    └─Sigmoid: 2-24                                    [1, 128, 88]              --
│    └─Linear: 2-25                                     [88, 128, 128]            32,896
=========================================================================================================
Total params: 5,516,574
Trainable params: 5,516,574
Non-trainable params: 0
Total mult-adds (M): 688.90
=========================================================================================================
Input size (MB): 0.20
Forward/backward pass size (MB): 3820.50
Params size (MB): 22.07
Estimated Total Size (MB): 3842.77
=========================================================================================================

Torch性能分析结果

通过torch.profiler进行性能分析的结果如下:

----------------------------  ------------  ------------  ------------  ------------  ------------  ------------  
                        Name    Self CPU %      Self CPU   CPU total %     CPU total  CPU time avg    # of Calls  
----------------------------  ------------  ------------  ------------  ------------  ------------  ------------  
             model_inference         9.38%     430.768ms       100.00%        4.591s        4.591s             1  
                aten::matmul         1.23%      56.495ms        72.63%        3.335s      35.856ms            93  
                aten::linear         0.03%       1.353ms        55.15%        2.532s      35.667ms            71  
                    aten::mm        53.39%        2.451s        53.39%        2.451s      34.527ms            71  
                   aten::bmm        14.66%     673.043ms        14.66%     673.265ms      30.603ms            22  
               aten::softmax         0.00%     144.000us         4.67%     214.293ms      19.481ms            11  
              aten::_softmax         4.66%     214.149ms         4.66%     214.149ms      19.468ms            11  
                 aten::clone         0.02%     866.000us         4.57%     209.970ms       4.117ms            51  
                 aten::copy_         4.54%     208.307ms         4.54%     208.307ms       3.858ms            54  
             aten::clamp_min         1.69%      77.715ms         3.38%     155.278ms       8.627ms            18  
----------------------------  ------------  ------------  ------------  ------------  ------------  ------------  
Self CPU time total: 4.591s

虽然我能从性能分析结果中看到瓶颈,但不确定能否大幅提升模型的推理速度,请问有什么可行的优化方案吗?


内容的提问来源于stack exchange,提问作者Phys

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 16:57:02