FastAPI+HuggingFace推理服务CPU内存持续上涨问题求助
解决FastAPI+HuggingFace图像Captioning服务内存泄漏问题
核心问题分析
服务每处理一次请求CPU内存就递增数MB,最终耗尽导致不可用,核心原因大概率是未显式释放推理过程中产生的临时资源,或Python垃圾回收(GC)未及时清理循环引用/残留对象。结合你的代码,重点排查以下几点:
针对性修复方案
1. 显式释放图片资源与临时对象
get_image加载的图片对象(如PIL Image)、BytesIO流若未主动释放,会持续占用内存。修改__predict方法,在每次迭代后强制清理资源:
import gc import torch def __predict(self, images): preds = [] for image in images: img_obj = None try: img_obj = get_image(image) pred_result = self.predictor(img_obj)[0]['generated_text'] preds.append(pred_result) except Exception as e: logging.error(f"Image URL error {image}: {str(e)}") preds.append('ERROR. Could not retrieve image.') finally: # 关闭并删除图片对象 if img_obj is not None: if hasattr(img_obj, 'close'): img_obj.close() del img_obj # 手动触发垃圾回收 gc.collect() # 清理GPU缓存(若使用GPU) if torch.cuda.is_available(): torch.cuda.empty_cache() return preds
2. 排查get_image函数的资源泄漏
确保get_image没有保留图片的冗余引用:
- 如果是通过
requests.get获取图片流,需显式关闭响应对象:def get_image(url): resp = requests.get(url, stream=True) resp.raise_for_status() img = Image.open(BytesIO(resp.content)) resp.close() # 显式关闭响应释放连接资源 return img - 避免在
get_image中创建全局/持久化的对象。
3. 优化HuggingFace Pipeline的内存使用
- 初始化pipeline时禁用缓存,减少临时数据积累:
self.predictor = pipeline('image-to-text', model=self.model, feature_extractor=self.extractor, tokenizer=self.tokenizer, device=device_num, cache=False) - 在
run方法末尾清理临时变量,辅助GC回收:def run(self, links: list): pred = measure_time(self.__predict, links) result_list = [{'url': link, 'prediction': result} for result, link in zip(pred['result'], links)] # 清理临时变量 del pred gc.collect() return {'result': result_list}
4. 内存泄漏定位工具
若以上修复仍未解决,用tracemalloc定位具体泄漏点:
import tracemalloc @app.post("/predict") async def image_captioning(predict: Predict): tracemalloc.start() data = predict.data result = model.run(data) snapshot = tracemalloc.take_snapshot() top_stats = snapshot.statistics('lineno') print("[Top 10 memory leak sources]") for stat in top_stats[:10]: print(stat) tracemalloc.stop() return result
额外优化建议
- 避免在循环中创建不必要的对象(如每次迭代生成新字典),尽量复用变量;
- 检查
measure_time函数,确保它未保留函数调用的结果引用; - 若服务流量较大,可采用多进程模式启动uvicorn(
uvicorn router:app --port 8000 --host 0.0.0.0 --workers 2),利用操作系统的进程回收机制缓解内存问题。
内容的提问来源于stack exchange,提问作者Guntram
相关产品推荐
相关产品推荐

