You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch Profiler导出栈文件为空问题求助

PyTorch Profiler导出栈信息为空的问题排查与解决

问题概述

复现PyTorch官方Profiler示例时,尝试导出模型前向传播的栈信息,但生成的栈文件为空;打印性能表格时也无法显示栈信息,调试发现EventList中每个事件的stacks属性均为空。

环境信息

  • Docker镜像:pytorch/pytorch:2.0.0-cuda11.7-cudnn8-runtime
  • torch版本:2.0.0
  • torchvision版本:0.15.0
  • Python版本:3.10.9

问题原因

  1. 默认采样模式的限制:torch.profiler.profile默认采用采样模式(采样频率1000Hz),单次模型前向传播的时间较短,采样器无法捕获足够的事件及对应的栈信息。
  2. 单次运行样本不足:仅执行一次前向传播,事件数量少,Profiler难以有效采集栈数据。
  3. CUDA栈信息捕获条件缺失:CUDA操作的栈信息依赖CPU与CUDA事件的关联记录,默认配置下未开启足够的关联参数,导致CUDA事件无法绑定对应栈信息。

解决方案

修改Profiler配置,改用追踪模式并增加模型运行次数,同时开启必要的关联参数:

import torch
from torch import profiler
from torchvision.models import resnet18

model = resnet18().cuda()
inputs = torch.rand(5, 3, 224, 224).cuda()

# 预热模型,确保CUDA kernel加载完成
model(inputs)

with profiler.profile(
    activities=[profiler.ProfilerActivity.CPU, profiler.ProfilerActivity.CUDA],
    with_stack=True,          # 开启栈追踪
    with_trace=True,          # 切换到追踪模式,记录所有操作事件
    record_shapes=True,       # 记录张量形状,增强事件关联性
    profile_memory=True       # 记录内存使用,辅助事件关联
) as p:
    # 多次运行模型,积累足够的事件样本
    for _ in range(10):
        model(inputs)

# 导出栈信息
p.export_stacks("/tmp/profiler/stacks_cpu.txt", "self_cpu_time_total")
p.export_stacks("/tmp/profiler/stacks_cuda.txt", "self_cuda_time_total")

# 打印带栈信息的性能表格
print(p.key_averages(group_by_stack_n=5).table(
    sort_by="self_cuda_time_total", row_limit=10))

关键修改说明

  • with_trace=True:从采样模式切换到追踪模式,Profiler会记录每一个操作事件,确保栈信息被完整捕获。
  • 多次运行模型:通过循环执行前向传播,让Profiler积累足够的事件样本,避免因单次运行时间过短导致的栈信息丢失。
  • record_shapes=True + profile_memory=True:增强CPU与CUDA事件的关联,确保CUDA操作能绑定到对应的CPU调用栈。

内容的提问来源于stack exchange,提问作者MufasaChan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 18:15:36