You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行Whisper后GPU内存满载,如何不重启Colab会话释放VRAM?

在Google Colab运行Whisper时GPU内存满载且无法释放的解决方法

问题场景

在Google Colab上运行Whisper的large-v1模型进行音频转录时,触发GPU内存不足错误,报错后GPU内存仍处于满载状态,需要无需重启会话的内存释放方案,避免已下载的音频数据丢失。


无需重启会话的GPU内存释放方案

1. 快速清空CUDA缓存

执行以下代码,释放PyTorch中未被引用的张量占用的GPU内存:

import torch
torch.cuda.empty_cache()

2. 强制垃圾回收+模型实例清理

如果仅清空缓存无效,手动删除模型对象并触发Python垃圾回收,彻底释放内存:

import gc
import torch

# 删除已加载的模型实例
del model
# 触发强制垃圾回收
gc.collect()
# 再次清空CUDA缓存
torch.cuda.empty_cache()

3. 内存占用优化(预防措施)

在加载Whisper模型时启用FP16精度,可大幅降低内存占用:

model = whisper.load_model("large-v1", device="cuda", dtype=torch.float16)

优化后的完整运行代码

以下代码加入了内存释放逻辑、FP16精度优化,并修复了原代码中的小问题:

!pip3 install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu117
!pip install git+https://github.com/openai/whisper.git 

import os
import torch
import gc
import whisper
from pathlib import Path
from whisper.utils import write_srt, write_txt, write_vtt
import pandas as pd

# 创建必要文件夹
if not os.path.exists("content"):
  os.mkdir("content")
if not os.path.exists("download"):
  os.mkdir("download")

def main():
    # 启用FP16精度加载模型,降低GPU内存占用
    model = whisper.load_model("large-v1", device="cuda", dtype=torch.float16)
    transcribe_name_begin = "oop"
    sub_folder_name = "/download/oop/"

    if not os.path.isdir(sub_folder_name):
        os.makedirs(sub_folder_name)
        
    _compression_ratio_threshold = 2.4
    for lectureId in range(142, 143):
        transcribePath = f"../content/{transcribe_name_begin}{lectureId}.mp3"
        try:
            result = model.transcribe(
                transcribePath,
                language="en",
                beam_size=9,
                initial_prompt="Welcome to the Software Engineering Courses channel.",
                best_of=9,
                verbose=True,
                temperature=0.0,
                compression_ratio_threshold=_compression_ratio_threshold
            )
            # 保存SRT字幕
            sub_name = f"{sub_folder_name}{transcribe_name_begin}{lectureId}.srt"
            with open(sub_name, "w", encoding="utf-8") as srt:
                write_srt(result["segments"], file=srt)

            # 保存多种格式转录结果
            writing_lut = {
                '.txt': write_txt,
                '.vtt': write_vtt,
                '.srt': write_srt,  # 修复原代码中srt对应write_txt的错误
            }

            output_type = "All"
            if output_type == "All":
                for suffix, write_func in writing_lut.items():
                    transcript_local_path = f"{sub_folder_name}{transcribe_name_begin}{lectureId}{suffix}"
                    with open(transcript_local_path, "w", encoding="utf-8") as f:
                        write_func(result["segments"], file=f)
                    print(f"**转录文件已创建: {transcript_local_path}**")
        except RuntimeError as e:
            if "out of memory" in str(e):
                print("GPU内存不足,正在释放内存...")
                # 清理当前批次的临时变量
                if 'result' in locals():
                    del result
                gc.collect()
                torch.cuda.empty_cache()
                print("已释放内存,建议降低beam_size或best_of参数后重试")
            else:
                raise e
    # 任务完成后彻底释放模型内存
    del model
    gc.collect()
    torch.cuda.empty_cache()
    print("所有任务完成,已释放GPU内存")

if __name__ == "__main__":
    main()

额外优化建议

  • 降低计算密集型参数:将beam_size和best_of从9调整为5,可显著减少内存占用,同时对转录质量影响极小
  • 长音频分段处理:将大型音频文件分割为10-30分钟的片段,逐个转录后合并结果,避免单批次内存过载
  • 循环内及时清理:在每次转录完成后,删除result等临时变量并触发垃圾回收,防止内存累积

内容的提问来源于stack exchange,提问作者Furkan Gözükara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 08:01:03