You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

torch.autograd.profiler添加record_function后CUDA时间异常及全步骤时间咨询

PyTorch Autograd Profiler 多头注意力分析问题解答

1. 为何添加不同的record_function后CUDA时间结果会变化?

CUDA操作是异步执行的,PyTorch默认不会等待CUDA任务完成就继续执行CPU代码。profiler.record_function在进入和退出代码块时,会自动触发CUDA流同步(确保该块内的CUDA操作全部完成):

  • 仅包裹softmax步骤时:profiler只会在softmax的record_function块结束时同步CUDA,此时QK矩阵乘法等前置CUDA操作可能还在后台运行,未被纳入统计,因此只显示softmax的CUDA时间。
  • 同时包裹QK乘法步骤时:profiler会在QK乘法的record_function块结束时同步CUDA,此时QK乘法的操作已完成,会被统计入总时间;另外,多次profiling的调度差异、测量误差也会导致时间数值有波动。

2. 是否仅会显示添加了record_function的步骤的CUDA时间?

不是。torch.autograd.profiler默认会自动记录PyTorch内置算子(如aten::matmul、aten::softmax)的CUDA时间,这些算子会以自身的名称出现在profiling结果中。你之前看到其他步骤CUDA时间为0,是因为异步执行导致的统计遗漏,而非未记录。自定义的record_function只是将其包裹范围内的所有子事件(包括内置算子)的时间聚合到自定义名称下,方便归类查看。

3. 如何获取所有步骤的CUDA运行时间?

要完整统计所有CUDA步骤的时间,需解决异步同步问题,并正确配置profiler:

关键步骤:

  • 强制CUDA同步:在profiling块结束后调用torch.cuda.synchronize(),确保所有CUDA操作都完成后再生成统计结果。
  • 启用调用栈记录:在profiler.profile中添加with_stack=True参数,能查看更详细的算子调用层级。
  • 查看内置算子事件:不要仅关注自定义的record_function事件,查看结果中所有aten::xxx开头的内置算子,这些就是注意力模块的各个底层步骤。
  • 按CUDA时间排序:统计结果时按self_cuda_time_total排序,能直观看到耗时最多的步骤。

示例代码:

import torch
import torch.autograd.profiler as profiler

def multihead_attention(queries, keys, values):
    qk = torch.matmul(queries, keys.transpose(-2, -1))
    scaled_qk = qk / torch.sqrt(torch.tensor(keys.size(-1), dtype=torch.float32, device=queries.device))
    attn_weights = torch.softmax(scaled_qk, dim=-1)
    output = torch.matmul(attn_weights, values)
    return output

# 初始化CUDA张量
queries = torch.randn(1, 8, 64, 64).cuda()
keys = torch.randn(1, 8, 64, 64).cuda()
values = torch.randn(1, 8, 64, 64).cuda()

with profiler.profile(
    use_cuda=True,
    record_shapes=True,
    profile_memory=True,
    with_stack=True  # 启用调用栈记录
) as prof:
    # 可选:用自定义事件包裹整个注意力模块
    with profiler.record_function("MULTIHEAD_ATTENTION_FULL"):
        a = multihead_attention(queries, keys, values)
    # 手动同步所有CUDA操作
    torch.cuda.synchronize()

# 按自身CUDA时间排序,显示所有事件(row_limit=-1表示不限制行数)
print(prof.key_averages().table(sort_by="self_cuda_time_total", row_limit=-1))

参考资源:

PyTorch官方的Autograd Profiler文档包含详细的使用指南和参数说明,可直接查阅官方文档了解更多细节。

4. 这些事件的含义是什么?在哪里可以找到相关文档?

  • 自定义事件:你通过record_function命名的事件(如SOFTMAX PASS),是对其包裹范围内所有子操作的时间聚合,含义由你定义。
  • 内置算子事件:以aten::开头的事件是PyTorch的底层ATen算子,比如aten::matmul对应矩阵乘法,aten::softmax对应Softmax运算,每个算子的功能可参考PyTorch官方的ATen算子文档。
  • profiler字段含义:输出表格中的self_cuda_time_total表示事件自身的CUDA耗时(不包含子事件),cuda_time_total表示事件及其所有子事件的总CUDA耗时,cpu_time_total则是CPU侧的耗时统计,这些字段的详细解释可在PyTorch Autograd Profiler的官方文档中找到。

内容的提问来源于stack exchange,提问作者TTXXAA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 17:41:02