如何获取ONNX模型各层的推理耗时?
获取ONNX模型各层推理耗时的方法
ONNX Runtime提供内置的性能分析功能,可直接采集各层(节点)的执行耗时,以下是具体实现步骤:
方法:启用内置性能分析(Profiling)
1. 配置会话并开启性能分析
创建InferenceSession时,通过SessionOptions开启profiling,并指定结果文件前缀:
import onnxruntime as ort import numpy as np import json # 配置会话选项,开启性能分析 sess_options = ort.SessionOptions() sess_options.enable_profiling = True # 设置性能分析结果文件的前缀,生成的文件格式为 profile_result.trace.json sess_options.profile_file_prefix = "profile_result" # 创建推理会话 ort_sess = ort.InferenceSession('model.onnx', sess_options)
2. 执行推理触发数据采集
构造输入并运行推理,此时会自动采集各节点的执行时间:
# 构造测试输入 onnx_input = np.random.normal(size=[1, 3, 224, 224]).astype(np.float32) ort_inputs = {ort_sess.get_inputs()[0].name: onnx_input} # 执行推理,触发性能数据采集 ort_outs = ort_sess.run(None, ort_inputs)
3. 导出并解析性能数据
调用end_profiling停止采集并获取结果文件路径,然后解析JSON文件提取各层耗时:
# 停止性能分析,获取结果文件路径 profile_path = ort_sess.end_profiling() # 解析JSON格式的性能数据 with open(profile_path, 'r') as f: profile_data = json.load(f) # 提取并打印各层(Node类型事件)的耗时 print("各层推理耗时统计:") for event in profile_data: # 过滤出节点执行的事件 if event['cat'] == 'Node': # dur字段单位为微秒,转换为毫秒 print(f"层名称: {event['name']}, 耗时(ms): {event['dur']/1000:.4f}")
说明
- 性能结果中的
dur字段单位为微秒,除以1000可转换为毫秒。 - 如果使用GPU推理(如CUDA provider),profiling结果也会包含GPU端的节点执行耗时。
- 生成的JSON文件还包含其他系统事件(如内存分配),通过
cat字段过滤Node类型即可得到各层的执行时间。
内容的提问来源于stack exchange,提问作者rufuss
相关产品推荐
相关产品推荐

