从C++向Python返回pybind11::array_t时程序崩溃的问题排查
问题:C++通过pybind11返回numpy数组时程序崩溃
问题重现
用户实现了如下pybind11绑定的C++库:
#include <vector> #include <cstdint> #include <pybind11/pybind11.h> #include <pybind11/numpy.h> namespace py = pybind11; struct Foo { using UintArr = py::array_t<std::uint64_t, py::array::c_style | py::array::forcecast>; UintArr get() const { std::vector<std::uint64_t> v = { 0, 1, 2, 3, 4, 5 }; return UintArr(v.size(), v.data()); } }; PYBIND11_MODULE(foo, m) { py::class_<Foo>(m, "Foo") .def(py::init<>()) .def("get", [](const Foo& t) { py::gil_scoped_release release; return t.get(); }); }
在Python中调用时程序崩溃:
import foo f = foo.Foo() f.get() # 此处崩溃
通过GDB分析核心文件,崩溃发生在导入numpy.core.multiarray时,调用栈如下:
#0 ... in PyImport_Import () #1 ... in PyImport_ImportModule () #2 ... in pybind11::module_::import (name=0x7f369e4c65de "numpy.core.multiarray") at /usr/include/pybind11/pybind11.h:1195 #3 ... in pybind11::detail::npy_api::lookup () at /usr/include/pybind11/numpy.h:264 #4 ... in pybind11::detail::npy_api::get () at /usr/include/pybind11/numpy.h:193 #5 ... in pybind11::detail::npy_format_descriptor<unsigned long, void>::dtype () at /usr/include/pybind11/numpy.h:1285 #6 ... in pybind11::dtype::of<unsigned long> () at /usr/include/pybind11/numpy.h:584 #7 ... in pybind11::array::array<unsigned long> (this=0x7ffed48a2548, shape=..., strides=..., ptr=0x560d3412b5a0, base=...) at /usr/include/pybind11/numpy.h:763 #8 ... in pybind11::array_t<unsigned long, 17>::array_t (this=0x7ffed48a2548, count=6, ptr=0x560d3412b5a0, base=...) at /usr/include/pybind11/numpy.h:1070 #9 ... in Foo::get (this=0x560d341285b0) at /home/steve.lorimer/src/python/example/pybind.cpp:15
用户已确认numpy.core.multiarray可在Python中正常导入:
import numpy.core.multiarray numpy.core.multiarray.__file__ # 输出:'/usr/local/lib/python3.10/dist-packages/numpy/core/multiarray.py'
问题原因
崩溃的核心原因是在释放全局解释器锁(GIL)后执行了依赖Python API的操作:
py::gil_scoped_release会释放GIL,允许其他Python线程运行,但此时当前线程不能调用任何Python API(包括导入模块、创建numpy数组等操作)。Foo::get()方法中创建py::array_t对象时,pybind11需要导入numpy.core.multiarray并调用其API构造numpy数组,这一操作必须在持有GIL的情况下进行。- 由于lambda中提前释放了GIL,后续调用
t.get()触发的numpy数组构造操作处于无GIL状态,直接导致崩溃。
解决方案
调整GIL的释放范围,仅在纯C++的耗时计算逻辑中释放GIL,确保创建numpy数组的操作始终在持有GIL的状态下执行。
修改后的绑定代码
PYBIND11_MODULE(foo, m) { py::class_<Foo>(m, "Foo") .def(py::init<>()) .def("get", [](const Foo& t) { std::vector<std::uint64_t> v; // 仅在纯C++计算阶段释放GIL { py::gil_scoped_release release; // 这里放置实际的耗时C++计算逻辑 v = { 0, 1, 2, 3, 4, 5 }; } // 创建numpy数组时必须持有GIL return Foo::UintArr(v.size(), v.data()); }); }
额外说明
如果Foo::get()中的逻辑本身依赖Python API(比如必须在函数内构造numpy数组),则不应在调用该方法前释放GIL,直接移除py::gil_scoped_release即可:
PYBIND11_MODULE(foo, m) { py::class_<Foo>(m, "Foo") .def(py::init<>()) .def("get", &Foo::get); }
内容的提问来源于stack exchange,提问作者Steve Lorimer
相关产品推荐
相关产品推荐

