如何在TensorFlow Serving中不重启服务添加新模型?
Absolutely! TensorFlow Serving is built to handle dynamic model additions without disrupting existing running services. Here are two straightforward, production-ready approaches to meet your needs:
Approach 1: File-based Model Repository with Auto-Refresh
This method leverages TensorFlow Serving's built-in ability to monitor changes to a model configuration file or directory.
Step 1: Set up your model repository structure
Organize your models in a directory where each model gets its own subfolder (with integer versioned subfolders inside, since TF Serving uses integer versions to track model updates):
models_root/ model_A/ 1/ saved_model.pb variables/ model_B/ 1/ ... model_C/ 1/ ...
Step 2: Create an initial model config file
Make a models.config YAML file that defines your initial models. Example content:
model_config_list { config { name: "model_A" base_path: "/models_root/model_A/" model_platform: "tensorflow" } config { name: "model_B" base_path: "/models_root/model_B/" model_platform: "tensorflow" } config { name: "model_C" base_path: "/models_root/model_C/" model_platform: "tensorflow" } }
Step 3: Start TF Serving with config polling
Launch the server with a flag to enable periodic checks for config file changes. This tells TF Serving to automatically reload the config (and load new models) without restarting:
tensorflow_model_server \ --model_config_file=/path/to/models.config \ --model_config_file_poll_wait_seconds=60 \ # Checks for updates every 60 seconds (adjust as needed) --port=8500 \ --rest_api_port=8501
Step 4: Add model_D dynamically
- Add your
model_Dfiles to the repository in the same structure:models_root/ model_D/ 1/ saved_model.pb variables/ - Update your
models.configto include the new model:# Add this block to the existing model_config_list config { name: "model_D" base_path: "/models_root/model_D/" model_platform: "tensorflow" } - Save the config file. TF Serving will detect the change within the
model_config_file_poll_wait_secondswindow and loadmodel_Din the background—no downtime formodel_A,model_B, ormodel_C.
Approach 2: Dynamic Model Addition via API
For immediate model loading (no waiting for config polling), use TF Serving's REST or gRPC API to send a reload request directly.
Example with REST API
Send a POST request to the TF Serving REST endpoint. Replace localhost:8501 with your server's address:
curl -X POST http://localhost:8501/v1/models/model_D \ -H "Content-Type: application/json" \ -d '{ "model_platform": "tensorflow", "base_path": "/models_root/model_D/" }'
This will trigger an immediate reload, and model_D will be available for inference right away—again, no impact on your running models.
Key Notes
- Always use integer version numbers for your model subfolders (TF Serving automatically picks the highest version for inference).
- When using the config file approach, avoid setting
model_config_file_poll_wait_secondsto a value that's too low (e.g., <10 seconds) as it can add unnecessary overhead. - Both methods work seamlessly with existing models; TF Serving handles multi-model inference in parallel without interruptions.
内容的提问来源于stack exchange,提问作者Rodrigo Laguna

