You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

以下PyTorch代码中网络推理时间的单位是什么?

问题解答

你代码里测量得到的推理时间单位是毫秒(ms)。

原因说明

代码中通过starter.elapsed_time(ender)获取的时间差,PyTorch官方定义该方法的返回值单位为毫秒:它返回两个CUDA Event之间的耗时,精度可达0.5微秒,但结果以毫秒为单位输出。

性能测量代码

#### 这段代码用于测量推理时间、FLOPs、参数量等性能指标
dummy_input = torch.randn(1, 3, 256, 256).cuda()

macs, params = profile(model, inputs=(dummy_input,), verbose=0)
macs, params = clever_format([macs, params], "%.3f")
name = "SegDepthWithTwoDecoders"
print("<" * 50, name)
print("Flops:", macs)
print("Parameters:", params)

starter, ender = torch.cuda.Event(enable_timing=True), torch.cuda.Event(
    enable_timing=True
)

repetitions = 300
timings = np.zeros((repetitions, 1))

# 预热运行10次,消除初始开销
for _ in range(10):
    _ = model(dummy_input)

# 性能测量阶段
with torch.no_grad():
    for rep in range(repetitions):
        starter.record()
        _ = model(dummy_input)
        ender.record()
        # 等待GPU同步,确保时间测量准确
        torch.cuda.synchronize()
        curr_time = starter.elapsed_time(ender)
        timings[rep] = curr_time
print("time :", np.average(timings))        
### 性能测量代码结束

代码输出

<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<< SegDepthWithTwoDecoders
Flops: 1.924G
Parameters: 3.986M
time : 12.67153577486674

这里的12.67代表单次推理平均耗时约12.67毫秒。

内容的提问来源于stack exchange,提问作者ashah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 12:20:26