Python的C++扩展内存泄漏问题排查求助
内存泄漏问题排查:Python C++扩展中的Numpy数组引用计数错误
我在编写Python模块的C++函数时遇到内存管理问题,每次调用该函数,Python解释器的内存占用都会持续增加。这个函数接收两个Numpy数组,生成新的Numpy数组作为输出。当从Python传入输出数组而非在函数内分配时,内存泄漏就消失了,我怀疑是引用计数处理错误导致返回的Numpy数组无法被垃圾回收,但调整引用计数后仍未解决。
附相关代码:
extern "C" PyObject *radec_to_xyz(PyObject *self, PyObject *args) { // Parse the input arguments PyArrayObject *ra_arrobj, *dec_arrobj; if (!PyArg_ParseTuple(args, "O!O!", &PyArray_Type, &ra_arrobj, &PyArray_Type, &dec_arrobj)) { PyErr_SetString(PyExc_TypeError, "invalid arguments, expected two numpy arrays"); return nullptr; } // skipping checks that would ensure: // dtype==float64, dim==1, len()>0 and equal for inputs, data or contiguous npy_intp size = PyArray_SIZE(ra_arrobj); // create the output numpy array with the same size and datatype PyObject *x_obj = PyArray_EMPTY(1, &size, NPY_FLOAT64, 0); if (!x_obj) return nullptr; Py_XINCREF(x_obj); // get pointers to the arrays double *ra_array = static_cast<double*>(PyArray_DATA(ra_arrobj)); double *dec_array = static_cast<double*>(PyArray_DATA(dec_arrobj)); double *x_array = static_cast<double*>(PyArray_DATA(reinterpret_cast<PyArrayObject*>(x_obj))); // compute the new coordinates for (npy_intp i = 0; i < size; ++i) { double cos_ra = cos(ra_array[i]); double cos_dec = cos(dec_array[i]); // compute final coordinates x_array[i] = cos_ra * cos_dec; } // return the arrays holding the new coordinates return Py_BuildValue("O", x_obj); }
问题根源
内存泄漏由两处错误的引用计数操作叠加导致:
PyArray_EMPTY创建数组后,已经返回一个引用计数为1的新对象,额外调用Py_XINCREF(x_obj)会将计数提升至2。Py_BuildValue("O", x_obj)会再次增加x_obj的引用计数,最终返回给Python时,对象的引用计数变为3。Python垃圾回收仅在计数降至0时触发,因此每次调用都会残留无法回收的引用,导致内存持续增长。
修复方案
移除多余的引用计数操作,直接返回PyArray_EMPTY创建的对象:
extern "C" PyObject *radec_to_xyz(PyObject *self, PyObject *args) { // Parse the input arguments PyArrayObject *ra_arrobj, *dec_arrobj; if (!PyArg_ParseTuple(args, "O!O!", &PyArray_Type, &ra_arrobj, &PyArray_Type, &dec_arrobj)) { PyErr_SetString(PyExc_TypeError, "invalid arguments, expected two numpy arrays"); return nullptr; } // skipping checks that would ensure: // dtype==float64, dim==1, len()>0 and equal for inputs, data or contiguous npy_intp size = PyArray_SIZE(ra_arrobj); // create the output numpy array with the same size and datatype PyObject *x_obj = PyArray_EMPTY(1, &size, NPY_FLOAT64, 0); if (!x_obj) return nullptr; // 移除多余的Py_XINCREF调用 // get pointers to the arrays double *ra_array = static_cast<double*>(PyArray_DATA(ra_arrobj)); double *dec_array = static_cast<double*>(PyArray_DATA(dec_arrobj)); double *x_array = static_cast<double*>(PyArray_DATA(reinterpret_cast<PyArrayObject*>(x_obj))); // compute the new coordinates for (npy_intp i = 0; i < size; ++i) { double cos_ra = cos(ra_array[i]); double cos_dec = cos(dec_array[i]); // compute final coordinates x_array[i] = cos_ra * cos_dec; } // 直接返回对象,无需通过Py_BuildValue额外增加引用计数 return x_obj; }
补充说明
PyArray_EMPTY返回的对象初始引用计数为1,直接返回给Python后,解释器会负责管理其生命周期:当Python变量不再引用该对象时,计数减1至0,触发垃圾回收释放内存。- 若因需求必须使用
Py_BuildValue,需先调用Py_DECREF(x_obj)抵消之前多余的Py_XINCREF,但直接返回对象是更简洁高效的方式。
内容的提问来源于stack exchange,提问作者van
相关产品推荐
相关产品推荐

