You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Google Colab中适配mBERT等模型到DST-MetaASSIST时遭遇CUDA内存不足错误

在Google Colab中适配mBERT等模型到DST-MetaASSIST时遭遇CUDA内存不足错误

我最近在尝试把mBERT、XLM-R和mT5这几个多语言预训练模型适配到DST-MetaASSIST(STAR)对话状态追踪项目里,但不管怎么调整,运行时总会碰到CUDA内存不足的错误。我已经试了几个优化方案,但还是没解决问题,想请教下大家有没有其他办法?

下面是我遇到的错误日志:

torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 20.00 MiB. GPU 0 has a total capacity of 14.75 GiB of which 9.06 MiB is free. Process 84806 has 14.74 GiB memory in use. Of the allocated memory 14.48 GiB is allocated by PyTorch, and 129.43 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.  See documentation for Memory Management

项目背景

我用到的DST-MetaASSIST项目可以通过以下方式获取:

### For DST-MetaASSIST
!git clone https://github.com/smartyfh/DST-MetaASSIST

待适配的多语言模型

我计划替换的三个模型的加载代码如下:

# 第一个模型:mBERT
from transformers import BertTokenizer, BertForSequenceClassification

# 加载mBERT分词器
tokenizer = BertTokenizer.from_pretrained('bert-base-multilingual-cased')
# 加载模型并设置正确的标签数量
model = BertForSequenceClassification.from_pretrained('bert-base-multilingual-cased', num_labels=num_labels)

# 第二个模型:XLM-R
from transformers import XLMRobertaTokenizer, XLMRobertaForSequenceClassification, Trainer, TrainingArguments

# 加载分词器和模型
tokenizer = XLMRobertaTokenizer.from_pretrained('xlm-roberta-base')
model = XLMRobertaForSequenceClassification.from_pretrained('xlm-roberta-base', num_labels=number_of_labels)

# 第三个模型:mT5
from transformers import T5Tokenizer, T5ForConditionalGeneration

# 加载mT5分词器和模型
tokenizer = T5Tokenizer.from_pretrained('google/mt5-small')
model = T5ForConditionalGeneration.from_pretrained('google/mt5-small')

已尝试的内存优化方案

为了解决内存问题,我已经做了这些调整:

  1. CUDA内存配置优化:设置环境变量PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,并调用torch.cuda.empty_cache()清空缓存
  2. 调小批次大小:把train_batch_size从原来的4/16降到2,meta_batch_size从2/8降到1
  3. 修改训练脚本:替换项目中train-S1.py的默认分词器为mBERT的分词器,调整相关参数配置。修改后的关键代码片段如下:
# 训练前的内存配置调整
import torch
import os

# 设置CUDA内存管理环境变量
os.environ['PYTORCH_CUDA_ALLOC_CONF'] = 'expandable_segments:True'
torch.cuda.empty_cache()  # 清空CUDA缓存
torch.backends.cudnn.benchmark = True

# 检查GPU可用性
if torch.cuda.is_available():
    device = torch.device("cuda")
    print("Using GPU:", torch.cuda.get_device_name(0))
else:
    device = torch.device("cpu")
    print("Using CPU")

# 调小批次大小
train_batch_size = 2  # 从4/16降低
meta_batch_size = 1    # 从2/8降低

# 启动训练命令
!python3 /content/DST-MetaASSIST/STAR/train-S1.py --data_dir data/mwz2.4 --save_dir output-meta24-S1/exp --train_batch_size 2 --meta_batch_size 1 --enc_lr 4e-5 --dec_lr 1e-4 --sw_lr 5e-5 --init_weight 0.5 --n_epochs 1 --do_train
# 脚本中替换分词器的核心代码
from transformers import BertTokenizer, BertForSequenceClassification

# 加载mBERT分词器
tokenizer = BertTokenizer.from_pretrained('bert-base-multilingual-cased')
# 或者通过命令行参数加载
tokenizer = BertTokenizer.from_pretrained(args.pretrained_model)

但即使做了这些调整,还是会触发内存不足的错误。有没有其他适合Colab环境的内存优化方法,能让我顺利跑起来这些多语言模型呢?


备注:内容来源于stack exchange,提问作者MarMarhoun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 15:09:30