使用MediaPipe与OpenCV训练手语检测模型时Python内存溢出原因排查
内存溢出原因及解决方案
核心原因
MediaPipe Hands实例重复创建且未释放
你的hand_crop函数每次调用都会新建mpHands.Hands(max_num_hands=10)实例,这个实例包含预加载的手部检测模型权重、计算图等大量内存资源。由于MediaPipe底层是C++实现,这些实例的内存不受Python垃圾回收器(GC)管理,每处理一张图片就新增一个未释放的实例,内存持续累积直至溢出。隐式引用导致内存无法回收
函数中生成的imgRGB、img_padding等数组,以及results对象,可能被MediaPipe实例隐式持有引用,导致Python GC无法识别这些对象已不再使用,进一步加剧内存占用。
解决方案
1. 复用MediaPipe Hands实例
不要在hand_crop内部每次创建实例,改为在数据集初始化时创建一次,所有图片处理共用同一个实例,避免重复加载模型资源。
2. 手动释放底层资源
在数据集销毁时调用hands.close()方法,手动释放MediaPipe底层的C++资源,避免内存泄漏。
3. 主动清理临时变量
在hand_crop处理完图片后,主动删除不再需要的临时变量,帮助GC更快回收内存。
修改后代码示例
手部裁剪函数修改
def hand_crop(img, hands, zoomout_ratio=2e-2): x = img.shape[1] y = img.shape[0] zoomout_x = x * zoomout_ratio zoomout_y = y * zoomout_ratio imgRGB = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) img_padding = cv2.copyMakeBorder(imgRGB, math.ceil(zoomout_y), math.ceil(zoomout_y), math.ceil(zoomout_x), math.ceil(zoomout_x), cv2.BORDER_CONSTANT, (0, 0, 0)) results = hands.process(img_padding) crop_result = [] if results.multi_hand_landmarks: for hlm in results.multi_hand_landmarks: x_max = int(max([i.x for i in hlm.landmark]) * (x + 2 * zoomout_x)) y_max = int(max([i.y for i in hlm.landmark]) * (y + 2 * zoomout_y)) x_min = int(min([i.x for i in hlm.landmark]) * (x + 2 * zoomout_x)) y_min = int(min([i.y for i in hlm.landmark]) * (y + 2 * zoomout_y)) width, height = x_max - x_min, y_max - y_min if width > height: y_min -= (((width - height) / 2) + zoomout_y) y_min = math.floor(y_min) y_max += (((width - height) / 2) + zoomout_y) y_max = math.floor(y_max) x_min -= zoomout_x x_min = math.floor(x_min) x_max += zoomout_x x_max = math.floor(x_max) else: x_min -= (((height - width) / 2) + zoomout_x) x_min = math.floor(x_min) x_max += (((height - width) / 2) + zoomout_x) x_max = math.floor(x_max) y_min -= zoomout_y y_min = math.floor(y_min) y_max += zoomout_y y_max = math.floor(y_max) crop = img_padding[y_min:y_max, x_min:x_max, :] cropBGR = cv2.cvtColor(crop, cv2.COLOR_RGB2BGR) crop_result.append(cropBGR) # 主动清理临时变量 del imgRGB, img_padding, results return crop_result
数据集类修改
class DatasetPlus(Dataset): def __init__(self, root_img, root_data, width, height, transform=None): self.root_img = root_img self.root_data = root_data self.width = width self.height = height self.transform = transform # 加载标签 self.labels = pd.read_csv(self.root_data) self.imgs = [image for image in sorted( os.listdir(self.root_img)) if image[-4:] == '.jpg'] self.len = len(self.imgs) # 初始化一次MediaPipe Hands实例,全局复用 self.mpHands = mp.solutions.hands self.hands = self.mpHands.Hands(max_num_hands=10) def __len__(self): return self.len def __getitem__(self, idx): img_name = self.imgs[idx] img_path = os.path.join(self.root_img, img_name) img = cv2.imread(img_path) print(img_name) # 传入已创建的hands实例 img = hand_crop(img, self.hands)[0] img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB).astype(np.float32) img = cv2.resize(img, (self.width, self.height), cv2.INTER_AREA) img = np.array(img) / 255.0 if self.transform is not None: img = self.transform(img) img_id = int(img_name[6:-4]) label = self.labels.where(self.labels['ID'] == img_id)['Label'].dropna().to_numpy()[0] if label == 3: label = 0 label = torch.tensor(label, dtype=torch.long) return img, label def __del__(self): # 手动释放MediaPipe资源 self.hands.close()
额外注意事项
如果你的DataLoader设置了num_workers > 0(多进程加载数据),每个worker会创建独立的DatasetPlus实例,此时每个worker都会初始化一个Hands实例。这种情况下可以考虑:
- 降低
num_workers数量,减少资源占用 - 在worker初始化时统一创建资源,退出时统一释放
内容的提问来源于stack exchange,提问作者yasamin tavakoli
相关产品推荐
相关产品推荐

