如何在Amazon SageMaker中加载已存至S3的训练后模型?
Great question! Let's break this down clearly for you:
1. 调用fit()后模型是否会保存到指定的S3存储桶?
Absolutely yes. When you set output_path=output_location while initializing the KMeans estimator, you’re explicitly telling SageMaker to store all training artifacts—including your trained model—in that exact S3 bucket path once fit() completes.
After training finishes successfully, SageMaker will automatically upload a compressed model archive (model.tar.gz) to your specified S3 location. The full path will look something like this:
s3://your-bucket/kmeans_highlevel_example/output/<your-training-job-name>/output/model.tar.gz
Here, <your-training-job-name> is either an auto-generated name or a custom one you set via the base_job_name parameter when creating the estimator. You can confirm this path by checking your S3 bucket directly, or by accessing kmeans.model_data right after fit() runs—it will return the exact S3 URL of the model archive.
2. 后续如何加载此模型?
There are two common approaches to load your trained KMeans model, depending on your use case:
方法一:通过训练后的Estimator直接加载(推荐)
If you still have access to the kmeans estimator object after training, you can use it to quickly create a model and deploy a prediction endpoint:
# Grab the model's S3 path from the estimator model_s3_path = kmeans.model_data # Initialize the KMeansModel from the S3 path from sagemaker.amazon.amazon_estimator import KMeansModel kmeans_model = KMeansModel(model_data=model_s3_path, role=role) # Deploy the model to a SageMaker endpoint predictor = kmeans_model.deploy(initial_instance_count=1, instance_type='ml.t2.medium') # Now you can use the predictor to make predictions # example_predictions = predictor.predict(your_input_data)
方法二:从已知的S3路径加载
If you’re working in a new script or don’t have the original estimator object, you can directly specify the model’s S3 URL to load it:
# Replace with your actual model S3 path model_s3_path = 's3://your-bucket/kmeans_highlevel_example/output/<your-training-job-name>/output/model.tar.gz' kmeans_model = KMeansModel(model_data=model_s3_path, role=role) predictor = kmeans_model.deploy(initial_instance_count=1, instance_type='ml.t2.medium')
额外:本地加载模型(无需部署)
If you want to test the model locally without deploying it to an endpoint, you can download and extract the model archive first:
- Use the AWS CLI to download the model file from S3:
aws s3 cp s3://your-bucket/kmeans_highlevel_example/output/<your-training-job-name>/output/model.tar.gz ./model.tar.gz
- Extract the compressed archive:
tar -xvzf model.tar.gz
- Load the model (SageMaker’s built-in KMeans uses a Scikit-learn model saved with joblib):
import joblib # Load the extracted model model = joblib.load('model.joblib') # Run local predictions # local_predictions = model.predict(your_local_dataset)
For production workflows, deploying to a SageMaker endpoint is generally recommended—it handles scaling, monitoring, and infrastructure management out of the box.
内容的提问来源于stack exchange,提问作者Borut Flis

