ComfyUI显存溢出后调用清理接口引发异常求助排查
显存溢出清理后触发CUDA异常的原因分析与解决思路
问题场景
在ComfyUI中执行图生视频、文生视频等显存密集型任务时,频繁遭遇OOM(显存溢出)。为了自动恢复,我在检测到OOM后调用ComfyUI-Easy-Use组件的cleangpu接口清理显存,代码逻辑如下:
if out of memory: SystemResourceManager.clear_queue() SystemResourceManager.clear_gpu() @classmethod def clear_gpu(cls): try: url = f"{cls.base_url}/api/easyuse/cleangpu" req = urllib.request.Request(url, method="POST") with urllib.request.urlopen(req, timeout=10) as response: if response.status == 200: import gc gc.collect() print("cleargpu success!") else: print(f"cleargpu failed: {response.status}") except Exception as e: print(f"cleargpu have error: {e}")
但执行后不仅OOM问题未解决,还触发了新的CUDA异常,完整错误日志如下:
!!! Exception during processing !!! Allocation on device 0 would exceed allowed memory. (out of memory) Currently allocated : 21.96 GiB Requested : 416.62 MiB Device limit : 23.50 GiB Free (according to CUDA): 31.69 MiB PyTorch limit (set by user-supplied memory fraction) : 17179869184.00 GiB Traceback (most recent call last): File "/root/ComfyUI/execution.py", line 323, in execute output_data, output_ui, has_subgraph = get_output_data(obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb) ...(省略中间栈信息) File "/usr/local/lib/python3.10/site-packages/diffusers/models/embeddings.py", line 440, in forward embeds = embeds + pos_embedding torch.cuda.OutOfMemoryError: Allocation on device 0 would exceed allowed memory. (out of memory) Got an OOM, unloading all loaded models. Exception in thread Thread-2 (prompt_worker): Traceback (most recent call last): ...(重复OOM栈信息) During handling of the above exception, another exception occurred: Traceback (most recent call last): File "/usr/local/lib/python3.10/threading.py", line 1009, in _bootstrap_inner self.run() ...(省略中间栈信息) File "/usr/local/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1158, in convert return t.to(device, dtype if t.is_floating_point() or t.is_complex() else None, non_blocking) RuntimeError: CUDA error: invalid argument CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1. Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
关键错误点提取
- 初始OOM:显存已用21.96GiB,剩余仅31.69MiB,无法满足416.62MiB的分配请求。
- 二次异常:ComfyUI自动执行
unload_all_models()卸载模型时,调用model.to(device_to)触发RuntimeError: CUDA error: invalid argument。
核心原因分析
- 清理时机与ComfyUI原生逻辑冲突:OOM发生后,ComfyUI会自动启动模型卸载流程(日志中
Got an OOM, unloading all loaded models.),此时用户代码同时调用外部cleangpu接口,导致双重内存操作,打乱了ComfyUI对模型状态的追踪,引发后续转移模型时的参数异常。 - 外部清理接口与ComfyUI内存管理不兼容:ComfyUI通过
model_management模块统一管理模型加载/卸载、显存分配,而cleangpu接口可能直接操作CUDA显存或独立处理模型状态,未与ComfyUI的内部状态同步,导致模型卸载时出现无效参数。 - CUDA上下文异常传递:OOM已经导致CUDA上下文处于异常状态,后续的模型转移操作(
model.to())会继承该异常状态,触发"invalid argument"错误——这类错误通常是因为之前的CUDA操作失败,导致后续调用的参数或环境无效。 - PyTorch内存配置错误:日志中
PyTorch limit (set by user-supplied memory fraction)显示为17179869184.00 GiB,这明显是配置错误(远超显卡实际显存),会导致PyTorch的内存限制逻辑失效,加剧OOM风险。
解决建议
- 移除自定义OOM清理逻辑:依赖ComfyUI原生的OOM处理机制即可,它会自动执行模型卸载和内存清理,不需要额外调用外部接口。
- 同步清理逻辑与ComfyUI执行流程:如果必须自定义清理,需要通过ComfyUI提供的钩子函数(如
pre_execute_cb、execution_block_cb)在任务执行间隙清理,避免在OOM异常抛出后并发操作。 - 修复PyTorch内存配置:检查启动参数或配置文件,将
--gpu-memory-fraction设置为合理值(比如0.9),确保PyTorch的内存限制符合显卡实际容量。 - 优化视频生成显存占用:
- 降低视频分辨率、帧率或时长
- 启用模型分片(model splitting)或CPU offload
- 使用fp16精度加载模型
内容的提问来源于stack exchange,提问作者Yiyabo
相关产品推荐
相关产品推荐

