You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在TensorFlow Serving中不重启服务添加新模型?

Absolutely! TensorFlow Serving is built to handle dynamic model additions without disrupting existing running services. Here are two straightforward, production-ready approaches to meet your needs:

Approach 1: File-based Model Repository with Auto-Refresh

This method leverages TensorFlow Serving's built-in ability to monitor changes to a model configuration file or directory.

Step 1: Set up your model repository structure

Organize your models in a directory where each model gets its own subfolder (with integer versioned subfolders inside, since TF Serving uses integer versions to track model updates):

models_root/
  model_A/
    1/
      saved_model.pb
      variables/
  model_B/
    1/
      ...
  model_C/
    1/
      ...

Step 2: Create an initial model config file

Make a models.config YAML file that defines your initial models. Example content:

model_config_list {
  config {
    name: "model_A"
    base_path: "/models_root/model_A/"
    model_platform: "tensorflow"
  }
  config {
    name: "model_B"
    base_path: "/models_root/model_B/"
    model_platform: "tensorflow"
  }
  config {
    name: "model_C"
    base_path: "/models_root/model_C/"
    model_platform: "tensorflow"
  }
}

Step 3: Start TF Serving with config polling

Launch the server with a flag to enable periodic checks for config file changes. This tells TF Serving to automatically reload the config (and load new models) without restarting:

tensorflow_model_server \
  --model_config_file=/path/to/models.config \
  --model_config_file_poll_wait_seconds=60 \  # Checks for updates every 60 seconds (adjust as needed)
  --port=8500 \
  --rest_api_port=8501

Step 4: Add model_D dynamically

  1. Add your model_D files to the repository in the same structure:
    models_root/
      model_D/
        1/
          saved_model.pb
          variables/
    
  2. Update your models.config to include the new model:
    # Add this block to the existing model_config_list
    config {
      name: "model_D"
      base_path: "/models_root/model_D/"
      model_platform: "tensorflow"
    }
    
  3. Save the config file. TF Serving will detect the change within the model_config_file_poll_wait_seconds window and load model_D in the background—no downtime for model_A, model_B, or model_C.

Approach 2: Dynamic Model Addition via API

For immediate model loading (no waiting for config polling), use TF Serving's REST or gRPC API to send a reload request directly.

Example with REST API

Send a POST request to the TF Serving REST endpoint. Replace localhost:8501 with your server's address:

curl -X POST http://localhost:8501/v1/models/model_D \
  -H "Content-Type: application/json" \
  -d '{
    "model_platform": "tensorflow",
    "base_path": "/models_root/model_D/"
  }'

This will trigger an immediate reload, and model_D will be available for inference right away—again, no impact on your running models.

Key Notes

  • Always use integer version numbers for your model subfolders (TF Serving automatically picks the highest version for inference).
  • When using the config file approach, avoid setting model_config_file_poll_wait_seconds to a value that's too low (e.g., <10 seconds) as it can add unnecessary overhead.
  • Both methods work seamlessly with existing models; TF Serving handles multi-model inference in parallel without interruptions.

内容的提问来源于stack exchange,提问作者Rodrigo Laguna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:54:38