You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取ONNX模型各层的推理耗时?

获取ONNX模型各层推理耗时的方法

ONNX Runtime提供内置的性能分析功能,可直接采集各层(节点)的执行耗时,以下是具体实现步骤:

方法:启用内置性能分析(Profiling)

1. 配置会话并开启性能分析

创建InferenceSession时,通过SessionOptions开启profiling,并指定结果文件前缀:

import onnxruntime as ort
import numpy as np
import json

# 配置会话选项,开启性能分析
sess_options = ort.SessionOptions()
sess_options.enable_profiling = True
# 设置性能分析结果文件的前缀,生成的文件格式为 profile_result.trace.json
sess_options.profile_file_prefix = "profile_result"

# 创建推理会话
ort_sess = ort.InferenceSession('model.onnx', sess_options)

2. 执行推理触发数据采集

构造输入并运行推理,此时会自动采集各节点的执行时间:

# 构造测试输入
onnx_input = np.random.normal(size=[1, 3, 224, 224]).astype(np.float32)
ort_inputs = {ort_sess.get_inputs()[0].name: onnx_input}

# 执行推理,触发性能数据采集
ort_outs = ort_sess.run(None, ort_inputs)

3. 导出并解析性能数据

调用end_profiling停止采集并获取结果文件路径,然后解析JSON文件提取各层耗时:

# 停止性能分析,获取结果文件路径
profile_path = ort_sess.end_profiling()

# 解析JSON格式的性能数据
with open(profile_path, 'r') as f:
    profile_data = json.load(f)

# 提取并打印各层(Node类型事件)的耗时
print("各层推理耗时统计:")
for event in profile_data:
    # 过滤出节点执行的事件
    if event['cat'] == 'Node':
        # dur字段单位为微秒,转换为毫秒
        print(f"层名称: {event['name']}, 耗时(ms): {event['dur']/1000:.4f}")

说明

  • 性能结果中的dur字段单位为微秒,除以1000可转换为毫秒。
  • 如果使用GPU推理(如CUDA provider),profiling结果也会包含GPU端的节点执行耗时。
  • 生成的JSON文件还包含其他系统事件(如内存分配),通过cat字段过滤Node类型即可得到各层的执行时间。

内容的提问来源于stack exchange,提问作者rufuss

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 20:22:06