AWS Glue Job连Google BigQuery遇iam.getAccessToken权限拒绝
问题:AWS Glue Job通过Google工作负载身份联合连接BigQuery时遭遇403权限拒绝
需求为从MySQL数据库提取数据并写入Google BigQuery,原服务账号方案可正常运行,替换为Google工作负载身份联合后出现权限错误。
1. 现有Glue脚本
import os import sys import json import boto3 from google.auth.transport.requests import Request from google.auth import impersonated_credentials from google.auth.external_account import Credentials as ExternalAccountCredentials from google.auth import load_credentials_from_file from google.cloud import bigquery from awsglue.transforms import * from awsglue.utils import getResolvedOptions from pyspark.context import SparkContext from awsglue.context import GlueContext from awsglue.job import Job from awsglue.dynamicframe import DynamicFrame import logging # Set up logging logging.basicConfig(level=logging.INFO) logger = logging.getLogger(__name__) # AWS Glue Job Arguments args = getResolvedOptions(sys.argv, ['JOB_NAME']) sc = SparkContext() glueContext = GlueContext(sc) spark = glueContext.spark_session job = Job(glueContext) job.init(args['JOB_NAME'], args) # AWS S3 setup for credentials bucket_name = 'aws-glue-assets-123123123-eu-central-1' credentials_file_key = 'credentials/service-account-impersonation.json' local_credentials_path = '/tmp/service-account-impersonation.json' project_id = 'nonprod_project' # Function to download credentials file from S3 def download_credentials(): logger.info(f"Downloading credentials from bucket: {bucket_name}, key: {credentials_file_key}") s3 = boto3.client('s3') try: s3.head_object(Bucket=bucket_name, Key=credentials_file_key) s3.download_file(bucket_name, credentials_file_key, local_credentials_path) logger.info(f"Credentials file downloaded to {local_credentials_path}") except boto3.exceptions.S3UploadFailedError as e: logger.error(f"Error downloading credentials file: {e}") sys.exit(1) # Download credentials download_credentials() # Load external account credentials credentials, project = load_credentials_from_file(local_credentials_path) # Specify the target service account to impersonate target_principal = 'myapp@-nonprod.iam.gserviceaccount.com' target_scopes = ['https://www.googleapis.com/auth/cloud-platform'] # Create impersonated credentials impersonated_creds = impersonated_credentials.Credentials( source_credentials=credentials, target_principal=target_principal, target_scopes=target_scopes ) # Refresh the impersonated credentials request = Request() impersonated_creds.refresh(request) # Initialize the BigQuery client client = bigquery.Client(credentials=impersonated_creds, project=project) # Define your DynamicFrames def create_dynamic_frame(table_name): return glueContext.create_dynamic_frame.from_options( connection_type="mysql", connection_options={ "useConnectionProperties": "true", "dbtable": table_name, "connectionName": "myapp nonprod DB", }, transformation_ctx=f"{table_name}_dynamic_frame" ) tables = ['changes', 'pull_requests', 'event_change', 'team_members'] dynamic_frames = {table: create_dynamic_frame(table) for table in tables} # Drop duplicates def drop_duplicates(frame): return DynamicFrame.fromDF(frame.toDF().dropDuplicates(), glueContext, f"{frame.transformation_ctx}_drop_duplicates") dynamic_frames = {table: drop_duplicates(df) for table, df in dynamic_frames.items()} # Drop fields for specific tables drop_fields_frames = { 'pull_requests': DropFields.apply(frame=dynamic_frames['pull_requests'], paths=["organization_name", "comment_id"], transformation_ctx="pull_requests_drop_fields") } # Function to load data to BigQuery def load_to_bigquery(df, table_id): pandas_df = df.toDF().toPandas() job_config = bigquery.LoadJobConfig(write_disposition=bigquery.WriteDisposition.WRITE_APPEND) load_job = client.load_table_from_dataframe(pandas_df, table_id, job_config=job_config) load_job.result() # Wait for the job to complete logger.info(f"Loaded {load_job.output_rows} rows into {table_id}.") # Write data to BigQuery for table, df in dynamic_frames.items(): table_id = f'change_tracker.{table}' load_to_bigquery(df, table_id) for table, df in drop_fields_frames.items(): table_id = f'change_tracker.{table}' load_to_bigquery(df, table_id) job.commit()
2. Google工作负载身份联合凭证配置
{ "type": "external_account", "audience": "//iam.googleapis.com/projects/123455678/locations/global/workloadIdentityPools/mynonprodpool/providers/aws-devendi-nonprod", "subject_token_type": "urn:ietf:params:aws:token-type:aws4_request", "service_account_impersonation_url": "https://iamcredentials.googleapis.com/v1/projects/-/serviceAccounts/myapp@nonprod.iam.gserviceaccount.com:generateAccessToken", "token_url": "https://sts.googleapis.com/v1/token", "credential_source": { "environment_id": "aws1", "region_url": "http://169.254.169.254/latest/meta-data/placement/availability-zone", "url": "http://169.254.169.254/latest/meta-data/iam/security-credentials", "regional_cred_verification_url": "https://sts.{region}.amazonaws.com?Action=GetCallerIdentity&Version=2011-06-15" } }
3. 错误信息
Error Category: UNCLASSIFIED_ERROR; RefreshError: ('Unable to acquire impersonated credentials', '{ "error": { "code": 403, "message": "Permission \'iam.serviceAccounts.getAccessToken\' denied on resource (or it may not exist).", "status": "PERMISSION_DENIED", "details": [ { "@type": "type.googleapis.com/google.rpc.ErrorInfo", "reason": "IAM_PERMISSION_DENIED", "domain": "iam.googleapis.com", "metadata": { "permission": "iam.serviceAccounts.getAccessToken" } } ] } }\n')
4. AWS IAM角色信任策略
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "Service": "glue.amazonaws.com" }, "Action": "sts:AssumeRole" }, { "Effect": "Allow", "Principal": { "Federated": "arn:aws:iam::123123123:oidc-provider/oidc.googleapis.com" }, "Action": "sts:AssumeRoleWithWebIdentity", "Condition": { "StringEquals": { "oidc.googleapis.com:aud": "aws-devendi-nonprod" } } } ] }
问题排查与解决方案
核心原因
通过工作负载身份联合认证后的身份,没有被授予impersonate目标Google服务账号的权限,导致无法调用iam.serviceAccounts.getAccessToken接口生成凭证。
修复步骤
给工作负载身份池中的AWS身份授予服务账号 impersonation 权限
- 登录Google Cloud Console,找到目标服务账号
myapp@nonprod.iam.gserviceaccount.com - 进入服务账号的权限页面,点击添加权限
- 在主体中填入:
principal://iam.googleapis.com/projects/[GCP_PROJECT_ID]/locations/global/workloadIdentityPools/mynonprodpool/subject/[AWS_GLUE_ROLE_ARN]- 替换
[GCP_PROJECT_ID]为你的GCP项目ID(如123455678) - 替换
[AWS_GLUE_ROLE_ARN]为Glue Job使用的IAM角色完整ARN(比如arn:aws:iam::123123123:role/your-glue-job-execution-role)
- 替换
- 在角色中搜索并选择
Service Account Token Creator角色(该角色包含所需的iam.serviceAccounts.getAccessToken权限) - 保存配置
- 登录Google Cloud Console,找到目标服务账号
修正脚本中的服务账号邮箱错误
脚本中target_principal字段多了一个多余的减号,需从myapp@-nonprod.iam.gserviceaccount.com修正为myapp@nonprod.iam.gserviceaccount.com,与凭证配置中的服务账号保持一致。验证AWS IAM角色的S3访问权限
确保Glue Job使用的IAM角色拥有s3:GetObject权限,允许访问S3桶aws-glue-assets-123123123-eu-central-1下的credentials/service-account-impersonation.json文件。
内容的提问来源于stack exchange,提问作者Marko Vasic
相关产品推荐
相关产品推荐

