You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google Colab中处理图像以解决内存不足问题

解决Colab内存不足问题:代码优化+内存扩容方案

一、代码层面优化(核心解决思路)

你的代码一次性加载所有图像特征和标签到内存,这是内存溢出的主要原因。以下是针对性优化方案:

1. 改用生成器逐批返回数据

放弃一次性存储所有features和labels,用yield逐批生成数据,训练时直接迭代生成器,彻底避免内存堆积。

修改后的create_dataset示例:

import gc
import numpy as np

def create_dataset_generator(data_df, SEQUENCE_LENGTH):
    video_paths = data_df["video_name"].unique()
    for video in video_paths:
        directory = f"images/{video}"
        frames = os.listdir(directory)
        # 提前按id_ped分组,避免重复切片浪费内存
        grouped_df = data_df[data_df.video_name == video].groupby("id_ped")
        for id_ped, ped in grouped_df:
            ped_len = len(ped)
            skip_frames_window = ped_len // SEQUENCE_LENGTH
            # 直接用numpy数组存储列数据,省去tolist()的内存开销
            ped_frames = ped["frame"].values
            ped_x1 = ped["x1"].values
            ped_y1 = ped["y1"].values
            ped_x2 = ped["x2"].values
            ped_y2 = ped["y2"].values
            ped_out1 = ped["out1"].values
            ped_out2 = ped["out2"].values
            ped_out3 = ped["out3"].values
            ped_out4 = ped["out4"].values
            
            for j in range(skip_frames_window):
                start_idx = SEQUENCE_LENGTH * j
                end_idx = SEQUENCE_LENGTH * (j + 1)
                pedestrian_list = []
                # 只遍历需要的帧区间,避免全量循环
                for i in range(start_idx, end_idx):
                    if i >= ped_len:
                        break
                    fr = ped_frames[i]
                    frame = frames[fr]
                    frame_path = f"{directory}/{frame}"
                    result = pedestrian_extraction(frame_path, ped_x1[i], ped_y1[i], ped_x2[i], ped_y2[i])
                    # 转换为float32,内存占用直接减半(默认float64)
                    result = result.astype(np.float32)
                    pedestrian_list.append(result)
                
                # 逐批返回数据
                yield np.array(pedestrian_list), [ped_out1[i], ped_out2[i], ped_out3[i], ped_out4[i]]
                # 手动清理临时变量,强制回收内存
                del pedestrian_list
                gc.collect()

训练时直接迭代生成器:

# 示例:训练循环中逐批获取数据
train_generator = create_dataset_generator(train_df, SEQUENCE_LENGTH)
for batch_features, batch_labels in train_generator:
    model.train_on_batch(batch_features, batch_labels)

2. 减少不必要的内存开销

  • 删掉所有.values.tolist()调用,直接用ped["frame"].values[i]访问数据,避免重复创建列表;
  • 图像特征统一转换为float32类型,大幅降低内存占用;
  • 每次循环后手动删除临时变量并调用gc.collect(),强制释放闲置内存。

3. 分块保存数据到磁盘

如果必须提前预处理数据,用h5py分块存储到磁盘,训练时按需加载:

import h5py

def save_dataset_to_disk(data_df, save_path, SEQUENCE_LENGTH, image_shape):
    with h5py.File(save_path, "w") as f:
        # 创建可动态扩展的数据集
        features_ds = f.create_dataset("features", shape=(0, SEQUENCE_LENGTH, *image_shape), maxshape=(None, SEQUENCE_LENGTH, *image_shape), dtype=np.float32)
        labels1_ds = f.create_dataset("labels1", shape=(0,), maxshape=(None,), dtype=np.int32)
        labels2_ds = f.create_dataset("labels2", shape=(0,), maxshape=(None,), dtype=np.int32)
        labels3_ds = f.create_dataset("labels3", shape=(0,), maxshape=(None,), dtype=np.int32)
        labels4_ds = f.create_dataset("labels4", shape=(0,), maxshape=(None,), dtype=np.int32)
        
        generator = create_dataset_generator(data_df, SEQUENCE_LENGTH)
        for batch_feat, batch_labels in generator:
            current_len = features_ds.shape[0]
            # 扩展数据集并写入数据
            features_ds.resize(current_len + 1, axis=0)
            features_ds[current_len] = batch_feat
            labels1_ds[current_len] = batch_labels[0]
            labels2_ds[current_len] = batch_labels[1]
            labels3_ds[current_len] = batch_labels[2]
            labels4_ds[current_len] = batch_labels[3]

二、Colab内存扩容方法

1. 升级到Colab Pro/Pro+

  • Pro版提供52GB高内存虚拟机选项,Pro+拥有更高资源优先级和更长运行时间;
  • 点击Colab界面右上角“升级到Pro”即可购买,适合长期重度使用。

2. 免费版Colab获取更大内存

  • 断开当前连接后重新连接虚拟机,空闲时段大概率能分配到27GB内存的实例(默认12GB);
  • 关闭其他闲置的Colab笔记本,释放系统资源。

3. 主动清理Colab内存

  • 定期删除无用变量:del train_features, train_labels,再执行gc.collect();
  • 通过“代码执行程序”→“管理会话”,手动终止旧会话释放内存。

4. 连接本地运行时

如果本地电脑有足够内存,可将Colab连接到本地Python环境:

  • 点击Colab界面“连接”→“连接到本地运行时”;
  • 按照提示在本地终端运行指定命令,即可利用本地内存处理数据。

内容的提问来源于stack exchange,提问作者RACHID BEN ABDELMALEK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 08:15:30