You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastAPI+HuggingFace推理服务CPU内存持续上涨问题求助

解决FastAPI+HuggingFace图像Captioning服务内存泄漏问题

核心问题分析

服务每处理一次请求CPU内存就递增数MB,最终耗尽导致不可用,核心原因大概率是未显式释放推理过程中产生的临时资源,或Python垃圾回收(GC)未及时清理循环引用/残留对象。结合你的代码,重点排查以下几点:

针对性修复方案

1. 显式释放图片资源与临时对象

get_image加载的图片对象(如PIL Image)、BytesIO流若未主动释放,会持续占用内存。修改__predict方法,在每次迭代后强制清理资源:

import gc
import torch

def __predict(self, images):
    preds = []
    for image in images:
        img_obj = None
        try:
            img_obj = get_image(image)
            pred_result = self.predictor(img_obj)[0]['generated_text']
            preds.append(pred_result)
        except Exception as e:
            logging.error(f"Image URL error {image}: {str(e)}")
            preds.append('ERROR. Could not retrieve image.')
        finally:
            # 关闭并删除图片对象
            if img_obj is not None:
                if hasattr(img_obj, 'close'):
                    img_obj.close()
                del img_obj
            # 手动触发垃圾回收
            gc.collect()
            # 清理GPU缓存(若使用GPU)
            if torch.cuda.is_available():
                torch.cuda.empty_cache()
    return preds

2. 排查get_image函数的资源泄漏

确保get_image没有保留图片的冗余引用:

  • 如果是通过requests.get获取图片流,需显式关闭响应对象:
    def get_image(url):
        resp = requests.get(url, stream=True)
        resp.raise_for_status()
        img = Image.open(BytesIO(resp.content))
        resp.close()  # 显式关闭响应释放连接资源
        return img
    
  • 避免在get_image中创建全局/持久化的对象。

3. 优化HuggingFace Pipeline的内存使用

  • 初始化pipeline时禁用缓存,减少临时数据积累:
    self.predictor = pipeline('image-to-text', 
                              model=self.model,
                              feature_extractor=self.extractor,
                              tokenizer=self.tokenizer, 
                              device=device_num,
                              cache=False)
    
  • 在run方法末尾清理临时变量,辅助GC回收:
    def run(self, links: list):
        pred = measure_time(self.__predict, links)
        result_list = [{'url': link, 'prediction': result} for result, link in
                      zip(pred['result'], links)]
        # 清理临时变量
        del pred
        gc.collect()
        return {'result': result_list}
    

4. 内存泄漏定位工具

若以上修复仍未解决,用tracemalloc定位具体泄漏点:

import tracemalloc

@app.post("/predict")
async def image_captioning(predict: Predict):
    tracemalloc.start()
    data = predict.data
    result = model.run(data)
    snapshot = tracemalloc.take_snapshot()
    top_stats = snapshot.statistics('lineno')
    print("[Top 10 memory leak sources]")
    for stat in top_stats[:10]:
        print(stat)
    tracemalloc.stop()
    return result

额外优化建议

  • 避免在循环中创建不必要的对象(如每次迭代生成新字典),尽量复用变量;
  • 检查measure_time函数,确保它未保留函数调用的结果引用;
  • 若服务流量较大,可采用多进程模式启动uvicorn(uvicorn router:app --port 8000 --host 0.0.0.0 --workers 2),利用操作系统的进程回收机制缓解内存问题。

内容的提问来源于stack exchange,提问作者Guntram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 07:16:22