如何避免调用Amazon SageMaker的fit方法时出现NoCredentialsError
Hey there! Let's break down how to fix that NoCredentialsError you're hitting when calling pca.fit()—it’s confusing since your S3 access works fine, but SageMaker’s estimator handles credentials a bit differently under the hood. Here are the most likely fixes, tailored to your code:
1. First, Check Your SageMaker Execution Role Permissions
The most common culprit here is that your AmazonSageMaker-ExecutionRole-SOMEVALUE doesn’t have the right permissions to access your S3 buckets. When SageMaker runs a training job, it uses this role to pull training data and push output artifacts—even if your local credentials work for S3, the role might not have access.
Make sure the role has at least these permissions:
s3:GetObjectfor your training data bucket (SOMEKINDOFBUCKETNAME)s3:PutObjectfor your output bucket (SOMEBUCKETNAME)
You can add a policy like this to the role via the IAM console:
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "s3:GetObject", "Resource": "arn:aws:s3:::SOMEKINDOFBUCKETNAME/sagemaker/pca/train/*" }, { "Effect": "Allow", "Action": ["s3:PutObject", "s3:GetObject"], "Resource": "arn:aws:s3:::SOMEBUCKETNAME/sagemaker/pca/output/*" } ] }
2. Ditch Hardcoded Credentials (They’re Risky & Can Cause Inconsistencies)
Hardcoding your AWS keys in the script isn’t a best practice, and it can lead to credential mismatches between your local S3 calls and the SageMaker estimator. Instead, let AWS handle credentials automatically:
- Use the AWS CLI to run
aws configureand set up your credentials locally - Or set the
AWS_ACCESS_KEY_IDandAWS_SECRET_ACCESS_KEYenvironment variables
Then simplify your session setup in code—you don’t need to manually create a boto3 session or SageMaker client. Let the SDK use the default credential chain:
# Replace your manual session/client setup with this sagemakerSession = sagemaker.Session() s3_resource = boto3.resource('s3') # Uses default credentials automatically
3. Verify Your Fit Input Format
While this isn’t a credential issue directly, using the TrainingInput class to define your training data can avoid path parsing quirks that might trigger misleading credential errors. Update your fit call like this:
from sagemaker.inputs import TrainingInput # Define your training input explicitly train_input = TrainingInput( s3_data=train_data, content_type='application/x-recordio-protobuf' ) pca.fit(inputs=train_input)
4. Double-Check Your Session Initialization (If You Must Use a Custom Session)
If you still need to use a custom boto3 session, make sure it’s properly passed to the SageMaker estimator and that it can access SageMaker services. Test it with a quick check:
session = boto3.session.Session( aws_access_key_id='YOUR_KEY', aws_secret_access_key='YOUR_SECRET', region_name='eu-west-1' ) # Test if the session can talk to SageMaker sm_client = session.client('sagemaker') print(sm_client.list_endpoints()) # Should return without errors if credentials are valid
Modified Code Example
Here’s your script with the key fixes applied:
import io import os import gzip import pickle import urllib.request import boto3 import sagemaker import sagemaker.amazon.common as smac from sagemaker.inputs import TrainingInput DOWNLOADED_FILENAME = 'C:/Users/Daan/PycharmProjects/downloads/mnist.pkl.gz' if not os.path.exists(DOWNLOADED_FILENAME): urllib.request.urlretrieve("http://deeplearning.net/data/mnist/mnist.pkl.gz", DOWNLOADED_FILENAME) with gzip.open(DOWNLOADED_FILENAME, 'rb') as f: train_set, valid_set, test_set = pickle.load(f, encoding='latin1') vectors = train_set[0].T buf = io.BytesIO() smac.write_numpy_to_dense_tensor(buf, vectors) buf.seek(0) key = 'recordio-pb-data' bucket_name = 'SOMEKINDOFBUCKETNAME' prefix = 'sagemaker/pca' path = os.path.join(prefix, 'train', key) # Use default credential chain sagemakerSession = sagemaker.Session() s3_resource = boto3.resource('s3') bucket = s3_resource.Bucket(bucket_name) current_bucket = bucket.Object(path) train_data = 's3://{}/{}/train/{}'.format(bucket_name, prefix, key) print('uploading training data location: {}'.format(train_data)) current_bucket.upload_fileobj(buf) output_location = 's3://{}/{}/output'.format('SOMEBUCKETNAME', prefix) print('training artifacts will be uploaded to: {}'.format(output_location)) region = 'eu-west-1' containers = { 'us-west-2': 'SOMELOCATION', 'us-east-1': 'SOMELOCATION', 'us-east-2': 'SOMELOCATION', 'eu-west-1': 'SOMELOCATION' } container = containers[region] role = 'AmazonSageMaker-ExecutionRole-SOMEVALUE' pca = sagemaker.estimator.Estimator( container, role, train_instance_count=1, train_instance_type='ml.c4.xlarge', output_path=output_location, sagemaker_session=sagemakerSession ) pca.set_hyperparameters( feature_dim=50000, num_components=10, subtract_mean=True, algorithm_mode='randomized', mini_batch_size=200 ) # Use explicit TrainingInput train_input = TrainingInput(s3_data=train_data, content_type='application/x-recordio-protobuf') pca.fit(inputs=train_input) print('END')
Start with checking the role permissions—this is almost always the fix for this specific error when S3 access works locally.
内容的提问来源于stack exchange,提问作者Daan

