You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用MediaPipe与OpenCV训练手语检测模型时Python内存溢出原因排查

内存溢出原因及解决方案

核心原因

  1. MediaPipe Hands实例重复创建且未释放
    你的hand_crop函数每次调用都会新建mpHands.Hands(max_num_hands=10)实例,这个实例包含预加载的手部检测模型权重、计算图等大量内存资源。由于MediaPipe底层是C++实现,这些实例的内存不受Python垃圾回收器(GC)管理,每处理一张图片就新增一个未释放的实例,内存持续累积直至溢出。

  2. 隐式引用导致内存无法回收
    函数中生成的imgRGB、img_padding等数组,以及results对象,可能被MediaPipe实例隐式持有引用,导致Python GC无法识别这些对象已不再使用,进一步加剧内存占用。

解决方案

1. 复用MediaPipe Hands实例

不要在hand_crop内部每次创建实例,改为在数据集初始化时创建一次,所有图片处理共用同一个实例,避免重复加载模型资源。

2. 手动释放底层资源

在数据集销毁时调用hands.close()方法,手动释放MediaPipe底层的C++资源,避免内存泄漏。

3. 主动清理临时变量

在hand_crop处理完图片后,主动删除不再需要的临时变量,帮助GC更快回收内存。

修改后代码示例

手部裁剪函数修改

def hand_crop(img, hands, zoomout_ratio=2e-2):
    x = img.shape[1]
    y = img.shape[0]

    zoomout_x = x * zoomout_ratio
    zoomout_y = y * zoomout_ratio

    imgRGB = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
    
    img_padding = cv2.copyMakeBorder(imgRGB, math.ceil(zoomout_y), math.ceil(zoomout_y), math.ceil(zoomout_x), math.ceil(zoomout_x), cv2.BORDER_CONSTANT, (0, 0, 0))

    results = hands.process(img_padding)

    crop_result = []

    if results.multi_hand_landmarks:
        for hlm in results.multi_hand_landmarks:
            x_max = int(max([i.x for i in hlm.landmark]) * (x + 2 * zoomout_x))   
            y_max = int(max([i.y for i in hlm.landmark]) * (y + 2 * zoomout_y))
            x_min = int(min([i.x for i in hlm.landmark]) * (x + 2 * zoomout_x))
            y_min = int(min([i.y for i in hlm.landmark]) * (y + 2 * zoomout_y))

            width, height = x_max - x_min, y_max - y_min

            if width > height:
                y_min -= (((width - height) / 2) + zoomout_y)
                y_min = math.floor(y_min)
                y_max += (((width - height) / 2) + zoomout_y)
                y_max = math.floor(y_max)
                x_min -= zoomout_x
                x_min = math.floor(x_min)
                x_max += zoomout_x
                x_max = math.floor(x_max)
            else:
                x_min -= (((height - width) / 2) + zoomout_x)
                x_min = math.floor(x_min)
                x_max += (((height - width) / 2) + zoomout_x)
                x_max = math.floor(x_max)
                y_min -= zoomout_y
                y_min = math.floor(y_min)
                y_max += zoomout_y
                y_max = math.floor(y_max)

            crop = img_padding[y_min:y_max, x_min:x_max, :]
            cropBGR = cv2.cvtColor(crop, cv2.COLOR_RGB2BGR)
            crop_result.append(cropBGR)
    
    # 主动清理临时变量
    del imgRGB, img_padding, results
    return crop_result

数据集类修改

class DatasetPlus(Dataset):
    def __init__(self, root_img, root_data, width, height, transform=None):
        self.root_img = root_img
        self.root_data = root_data
        self.width = width
        self.height = height
        self.transform = transform
        # 加载标签
        self.labels = pd.read_csv(self.root_data)
        self.imgs = [image for image in sorted(
            os.listdir(self.root_img)) if image[-4:] == '.jpg']
        self.len = len(self.imgs)
        # 初始化一次MediaPipe Hands实例,全局复用
        self.mpHands = mp.solutions.hands
        self.hands = self.mpHands.Hands(max_num_hands=10)

    def __len__(self):
        return self.len

    def __getitem__(self, idx):
        img_name = self.imgs[idx]
        img_path = os.path.join(self.root_img, img_name)
        img = cv2.imread(img_path)
        print(img_name)
        # 传入已创建的hands实例
        img = hand_crop(img, self.hands)[0]
        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB).astype(np.float32)
        img = cv2.resize(img, (self.width, self.height), cv2.INTER_AREA)
        img = np.array(img) / 255.0
        if self.transform is not None:
            img = self.transform(img)
        img_id = int(img_name[6:-4])
        label = self.labels.where(self.labels['ID'] == img_id)['Label'].dropna().to_numpy()[0]
        if label == 3: label = 0
        label = torch.tensor(label, dtype=torch.long)
        return img, label
    
    def __del__(self):
        # 手动释放MediaPipe资源
        self.hands.close()

额外注意事项

如果你的DataLoader设置了num_workers > 0(多进程加载数据),每个worker会创建独立的DatasetPlus实例,此时每个worker都会初始化一个Hands实例。这种情况下可以考虑:

  • 降低num_workers数量,减少资源占用
  • 在worker初始化时统一创建资源,退出时统一释放

内容的提问来源于stack exchange,提问作者yasamin tavakoli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 09:45:06