运行Whisper后GPU内存满载,如何不重启Colab会话释放VRAM?
在Google Colab运行Whisper时GPU内存满载且无法释放的解决方法
问题场景
在Google Colab上运行Whisper的large-v1模型进行音频转录时,触发GPU内存不足错误,报错后GPU内存仍处于满载状态,需要无需重启会话的内存释放方案,避免已下载的音频数据丢失。
无需重启会话的GPU内存释放方案
1. 快速清空CUDA缓存
执行以下代码,释放PyTorch中未被引用的张量占用的GPU内存:
import torch torch.cuda.empty_cache()
2. 强制垃圾回收+模型实例清理
如果仅清空缓存无效,手动删除模型对象并触发Python垃圾回收,彻底释放内存:
import gc import torch # 删除已加载的模型实例 del model # 触发强制垃圾回收 gc.collect() # 再次清空CUDA缓存 torch.cuda.empty_cache()
3. 内存占用优化(预防措施)
在加载Whisper模型时启用FP16精度,可大幅降低内存占用:
model = whisper.load_model("large-v1", device="cuda", dtype=torch.float16)
优化后的完整运行代码
以下代码加入了内存释放逻辑、FP16精度优化,并修复了原代码中的小问题:
!pip3 install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu117 !pip install git+https://github.com/openai/whisper.git import os import torch import gc import whisper from pathlib import Path from whisper.utils import write_srt, write_txt, write_vtt import pandas as pd # 创建必要文件夹 if not os.path.exists("content"): os.mkdir("content") if not os.path.exists("download"): os.mkdir("download") def main(): # 启用FP16精度加载模型,降低GPU内存占用 model = whisper.load_model("large-v1", device="cuda", dtype=torch.float16) transcribe_name_begin = "oop" sub_folder_name = "/download/oop/" if not os.path.isdir(sub_folder_name): os.makedirs(sub_folder_name) _compression_ratio_threshold = 2.4 for lectureId in range(142, 143): transcribePath = f"../content/{transcribe_name_begin}{lectureId}.mp3" try: result = model.transcribe( transcribePath, language="en", beam_size=9, initial_prompt="Welcome to the Software Engineering Courses channel.", best_of=9, verbose=True, temperature=0.0, compression_ratio_threshold=_compression_ratio_threshold ) # 保存SRT字幕 sub_name = f"{sub_folder_name}{transcribe_name_begin}{lectureId}.srt" with open(sub_name, "w", encoding="utf-8") as srt: write_srt(result["segments"], file=srt) # 保存多种格式转录结果 writing_lut = { '.txt': write_txt, '.vtt': write_vtt, '.srt': write_srt, # 修复原代码中srt对应write_txt的错误 } output_type = "All" if output_type == "All": for suffix, write_func in writing_lut.items(): transcript_local_path = f"{sub_folder_name}{transcribe_name_begin}{lectureId}{suffix}" with open(transcript_local_path, "w", encoding="utf-8") as f: write_func(result["segments"], file=f) print(f"**转录文件已创建: {transcript_local_path}**") except RuntimeError as e: if "out of memory" in str(e): print("GPU内存不足,正在释放内存...") # 清理当前批次的临时变量 if 'result' in locals(): del result gc.collect() torch.cuda.empty_cache() print("已释放内存,建议降低beam_size或best_of参数后重试") else: raise e # 任务完成后彻底释放模型内存 del model gc.collect() torch.cuda.empty_cache() print("所有任务完成,已释放GPU内存") if __name__ == "__main__": main()
额外优化建议
- 降低计算密集型参数:将
beam_size和best_of从9调整为5,可显著减少内存占用,同时对转录质量影响极小 - 长音频分段处理:将大型音频文件分割为10-30分钟的片段,逐个转录后合并结果,避免单批次内存过载
- 循环内及时清理:在每次转录完成后,删除
result等临时变量并触发垃圾回收,防止内存累积
内容的提问来源于stack exchange,提问作者Furkan Gözükara
相关产品推荐
相关产品推荐

