FastPitch模型二次/三次推理时出现CUDA OutOfMemory Error问题求助
FastPitch推理CUDA OutOfMemory问题排查与解决
核心问题:显存累积未释放
连续推理时,FastPitch的计算图、中间张量未被及时清理,导致显存逐步占用,最终在中等长度序列第二次推理、短序列第三次推理时触发OOM。
针对性解决方法:
- 每次推理后强制清理显存
在推理结束后立即执行显存清理操作,释放未被自动回收的张量:import torch # 清空CUDA缓存 torch.cuda.empty_cache() # 显式删除不再使用的中间变量(替换为实际变量名) del input_tensor, output_tensor - 用
torch.no_grad()包裹推理逻辑
推理阶段无需记录梯度,禁用梯度计算可避免额外显存占用:with torch.no_grad(): model_output = model(inputs) - 避免重复加载模型
确保模型仅初始化一次,重复加载会导致多份权重占用显存:# 正确做法:全局初始化一次模型 model = FastPitchModel.from_pretrained("your_model_path") model = model.cuda() def run_infer(input_data): with torch.no_grad(): return model(input_data) - 启用FP16混合精度推理
半精度计算可大幅降低显存消耗,FastPitch原生支持该模式:model = model.half().cuda() inputs = inputs.half().cuda() with torch.no_grad(): outputs = model(inputs) - 限制单批次序列长度/数量
即使短序列,多次推理的累积显存也可能超标,可适当减小单批次的序列数量或对过长序列做截断处理。
内容的提问来源于stack exchange,提问作者Victor Zheng
相关产品推荐
相关产品推荐

