torch.autograd.profiler添加record_function后CUDA时间异常及全步骤时间咨询
PyTorch Autograd Profiler 多头注意力分析问题解答
1. 为何添加不同的record_function后CUDA时间结果会变化?
CUDA操作是异步执行的,PyTorch默认不会等待CUDA任务完成就继续执行CPU代码。profiler.record_function在进入和退出代码块时,会自动触发CUDA流同步(确保该块内的CUDA操作全部完成):
- 仅包裹softmax步骤时:profiler只会在softmax的
record_function块结束时同步CUDA,此时QK矩阵乘法等前置CUDA操作可能还在后台运行,未被纳入统计,因此只显示softmax的CUDA时间。 - 同时包裹QK乘法步骤时:profiler会在QK乘法的
record_function块结束时同步CUDA,此时QK乘法的操作已完成,会被统计入总时间;另外,多次profiling的调度差异、测量误差也会导致时间数值有波动。
2. 是否仅会显示添加了record_function的步骤的CUDA时间?
不是。torch.autograd.profiler默认会自动记录PyTorch内置算子(如aten::matmul、aten::softmax)的CUDA时间,这些算子会以自身的名称出现在profiling结果中。你之前看到其他步骤CUDA时间为0,是因为异步执行导致的统计遗漏,而非未记录。自定义的record_function只是将其包裹范围内的所有子事件(包括内置算子)的时间聚合到自定义名称下,方便归类查看。
3. 如何获取所有步骤的CUDA运行时间?
要完整统计所有CUDA步骤的时间,需解决异步同步问题,并正确配置profiler:
关键步骤:
- 强制CUDA同步:在profiling块结束后调用
torch.cuda.synchronize(),确保所有CUDA操作都完成后再生成统计结果。 - 启用调用栈记录:在
profiler.profile中添加with_stack=True参数,能查看更详细的算子调用层级。 - 查看内置算子事件:不要仅关注自定义的
record_function事件,查看结果中所有aten::xxx开头的内置算子,这些就是注意力模块的各个底层步骤。 - 按CUDA时间排序:统计结果时按
self_cuda_time_total排序,能直观看到耗时最多的步骤。
示例代码:
import torch import torch.autograd.profiler as profiler def multihead_attention(queries, keys, values): qk = torch.matmul(queries, keys.transpose(-2, -1)) scaled_qk = qk / torch.sqrt(torch.tensor(keys.size(-1), dtype=torch.float32, device=queries.device)) attn_weights = torch.softmax(scaled_qk, dim=-1) output = torch.matmul(attn_weights, values) return output # 初始化CUDA张量 queries = torch.randn(1, 8, 64, 64).cuda() keys = torch.randn(1, 8, 64, 64).cuda() values = torch.randn(1, 8, 64, 64).cuda() with profiler.profile( use_cuda=True, record_shapes=True, profile_memory=True, with_stack=True # 启用调用栈记录 ) as prof: # 可选:用自定义事件包裹整个注意力模块 with profiler.record_function("MULTIHEAD_ATTENTION_FULL"): a = multihead_attention(queries, keys, values) # 手动同步所有CUDA操作 torch.cuda.synchronize() # 按自身CUDA时间排序,显示所有事件(row_limit=-1表示不限制行数) print(prof.key_averages().table(sort_by="self_cuda_time_total", row_limit=-1))
参考资源:
PyTorch官方的Autograd Profiler文档包含详细的使用指南和参数说明,可直接查阅官方文档了解更多细节。
4. 这些事件的含义是什么?在哪里可以找到相关文档?
- 自定义事件:你通过
record_function命名的事件(如SOFTMAX PASS),是对其包裹范围内所有子操作的时间聚合,含义由你定义。 - 内置算子事件:以
aten::开头的事件是PyTorch的底层ATen算子,比如aten::matmul对应矩阵乘法,aten::softmax对应Softmax运算,每个算子的功能可参考PyTorch官方的ATen算子文档。 - profiler字段含义:输出表格中的
self_cuda_time_total表示事件自身的CUDA耗时(不包含子事件),cuda_time_total表示事件及其所有子事件的总CUDA耗时,cpu_time_total则是CPU侧的耗时统计,这些字段的详细解释可在PyTorch Autograd Profiler的官方文档中找到。
内容的提问来源于stack exchange,提问作者TTXXAA
相关产品推荐
相关产品推荐

