You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Coqui TTS框架微调VITS模型时CUDA显存不足问题求助

Coqui TTS VITS微调CUDA内存不足解决方案

核心优化措施

  • 截断长音频样本:数据集内最大音频长度达389781采样点(按16kHz采样率计算约24秒),VITS对长音频的内存开销极高。在数据预处理阶段设置max_audio_len参数,将过长音频截断至200000采样点(约12.5秒),过滤或裁剪超出阈值的样本,可大幅降低单步训练的内存占用。
  • 启用混合精度训练:在Trainer初始化时添加precision=16参数,利用Colab GPU的FP16计算支持,能将内存占用降低约50%,同时不影响训练精度。
  • 调整梯度累积参数:即使batch size设为1,可设置accumulate_grad_batches=4,等效于batch size=4的训练效果,但内存占用仅为单batch的开销,平衡训练效率与内存压力。
  • 降低Batch Group Size:当前设置的Batch group size=256过大,会导致数据加载阶段预分配过多内存,建议降至64或32,减少内存碎片与预占用。
  • 关闭非必要监控:禁用TensorBoard实时日志、额外的指标计算等功能,避免这些组件占用GPU内存。

额外内存清理技巧

  • 重启Colab会话:GPU内存碎片可能由会话累积的临时变量导致,重启后重新运行所有步骤,可彻底清理残留内存。
  • 主动释放内存:在trainer.fit()执行前,除调用torch.cuda.empty_cache()外,手动删除不再需要的变量(如预处理后的临时数据),再执行缓存清理:
    del temp_data_var
    torch.cuda.empty_cache()
    
  • 优化模型加载:加载预训练VITS模型时,直接指定map_location='cuda',避免先加载到CPU再转移至GPU产生的临时内存占用:
    model = torch.load('pretrained_vits.pth', map_location='cuda')
    

问题背景

执行YouTube视频《Updated | Near-Automated Voice Cloning | Whisper STT + Coqui TTS | Fine Tune a VITS Model on Colab》对应的Colab Notebook,基于Coqui TTS框架微调VITS模型时,运行到trainer.fit()环节出现CUDA内存不足错误:

OutOfMemoryError: CUDA out of memory. Tried to allocate 12.00 MiB (GPU 0; 14.75 GiB total capacity; 13.41 GiB already allocated; 2.81 MiB free; 13.69 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.  See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF

已尝试的优化操作:

  • 添加torch.cuda.empty_cache()代码
  • 将batch size降至1
  • 设置%env 'PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512'
  • 使用Colab Pro账号

数据集详情:

> DataLoader initialization
| > Tokenizer:
    | > add_blank: True
    | > use_eos_bos: False
    | > use_phonemes: True
    | > phonemizer:
        | > phoneme language: en-us
        | > phoneme backend: espeak
| > Number of instances : 131
 | > Preprocessing samples
 | > Max text length: 280
 | > Min text length: 30
 | > Avg text length: 142.8320610687023
 | 
 | > Max audio length: 389781.0
 | > Min audio length: 103209.0
 | > Avg audio length: 225843.31297709924
 | > Num. instances discarded samples: 0
 | > Batch group size: 256.

内容的提问来源于stack exchange,提问作者Ghulam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 21:20:40