C++嵌入Python TensorFlow时模型加载、predict调用无限阻塞无报错
问题背景
构建NIST FRVT11提交库时,核心推理逻辑采用Python编写,需通过C++侧调用,已参考官方Python C API文档实现基础调用逻辑,但出现两类阻塞问题:
- 初始化阶段可正常完成Keras模型加载与单次预测,后续再次调用同一模型的
model.predict时无限耗时无响应 - 部分场景下调用
tf.keras.models.load_model阶段就直接无限阻塞,全程无任何错误输出
因NIST要求所有依赖需打包进提交库,已本地编译Python环境随包附带。
环境版本
- Python: 3.9.13
- TensorFlow: 2.8.0
- Keras: 2.8.0
现有实现代码
Python侧代码
import tensorflow as tf def get_model(model_path): embedding_model = tf.keras.models.load_model(model_path) return embedding_model def predict(model, img): out= model.predict(img) return out
C++侧代码
int call_python_code(const std::string &configDir) { this->configDir = configDir; // add directory to sys path std::string command_sys_path_to_append = "import sys, os\n" "sys.path.append(os.path.join(os.getcwd(), '" + this- >configDir + "'))\n"; PyRun_SimpleString(command_sys_path_to_append.c_str()); Py_Initialize(); PyObject *p_module_name = PyUnicode_DecodeFSDefault("fr"); // fr.py file this->p_module = PyImport_Import(p_module_name); // load the model from c++ PyObject *outer_tuple_2 = PyTuple_New(1); PyTuple_SetItem(outer_tuple_2, 0, PyUnicode_DecodeFSDefault(this->configDir.c_str())); PyObject *p_fucntion_model = PyObject_GetAttrString(p_module, "get_model"); this->p_embedding_model_object = PyObject_CallObject(p_fucntion_model, outer_tuple_2); // get prediction PyObject *p_function_predict= PyObject_GetAttrString(p_module, "predict"); PyObject *full_args = PyTuple_New(2); PyTuple_SetItem(full_args, 0, this->p_fucntion_model); PyTuple_SetItem(full_args, 1, <image array>); PyObject *outer_tuple = PyTuple_New(1); PyTuple_SetItem(outer_tuple, 0, full_args); PyObject *p_return_tuple = PyObject_CallObject(p_function_predict, outer_tuple); }
问题原因与修复方案
现有代码存在4处明确错误,是触发无限阻塞的核心原因:
- Python解释器初始化顺序错误
代码在调用Py_Initialize()启动解释器之前,就先调用了PyRun_SimpleString执行sys.path修改逻辑。解释器未初始化时调用Python C API属于未定义行为,随机出现加载阻塞、调用异常都是典型表现。
修复方式:将Py_Initialize()移到所有Python C API调用的最前端,确认解释器初始化完成后,再执行sys路径追加、模块导入等操作。 - 预测接口传参对象错误
模型加载完成后,实例对象保存在this->p_embedding_model_object中,但调用predict函数时,传入的第一个参数是this->p_fucntion_model(即get_model函数对象本身),并非加载完成的模型实例。首次调用能运行属于内存脏数据巧合,后续调用时类型不匹配会触发TensorFlow内部GIL死锁,直接无限卡住。
修复方式:构造predict参数元组时,第一个参数传入this->p_embedding_model_object。 - 参数元组嵌套错误
Python侧predict函数定义接收2个位置参数,但代码在构造完长度为2的参数元组后,又额外套了一层长度为1的外层元组,相当于只给predict传了1个元组类型的参数,参数数量不匹配会触发Python运行时异常。代码未做异常捕获时,会出现静默阻塞。
修复方式:移除多余的外层元组,直接将组装好的2元素参数元组传入PyObject_CallObject即可。 - 缺少GIL持有校验与异常检查
C++侧多线程场景下调用Python接口如果未持有GIL,或者调用后未检查异常状态,很容易触发死锁、无报错阻塞的问题。
修复方式:所有Python C API调用前确保当前线程持有GIL,每次调用PyObject_CallObject等核心接口后,通过PyErr_Occurred()检查异常状态,出现异常时及时获取traceback打印排查。
内容的提问来源于stack exchange,提问作者Nawal
相关产品推荐
相关产品推荐

