You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Glue Job连Google BigQuery遇iam.getAccessToken权限拒绝

问题:AWS Glue Job通过Google工作负载身份联合连接BigQuery时遭遇403权限拒绝

需求为从MySQL数据库提取数据并写入Google BigQuery,原服务账号方案可正常运行,替换为Google工作负载身份联合后出现权限错误。

1. 现有Glue脚本

import os
import sys
import json
import boto3
from google.auth.transport.requests import Request
from google.auth import impersonated_credentials
from google.auth.external_account import Credentials as ExternalAccountCredentials
from google.auth import load_credentials_from_file
from google.cloud import bigquery
from awsglue.transforms import *
from awsglue.utils import getResolvedOptions
from pyspark.context import SparkContext
from awsglue.context import GlueContext
from awsglue.job import Job
from awsglue.dynamicframe import DynamicFrame
import logging

# Set up logging
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

# AWS Glue Job Arguments
args = getResolvedOptions(sys.argv, ['JOB_NAME'])
sc = SparkContext()
glueContext = GlueContext(sc)
spark = glueContext.spark_session
job = Job(glueContext)
job.init(args['JOB_NAME'], args)

# AWS S3 setup for credentials
bucket_name = 'aws-glue-assets-123123123-eu-central-1'
credentials_file_key = 'credentials/service-account-impersonation.json'
local_credentials_path = '/tmp/service-account-impersonation.json'
project_id = 'nonprod_project'

# Function to download credentials file from S3
def download_credentials():
    logger.info(f"Downloading credentials from bucket: {bucket_name}, key: {credentials_file_key}")
    s3 = boto3.client('s3')
    try:
        s3.head_object(Bucket=bucket_name, Key=credentials_file_key)
        s3.download_file(bucket_name, credentials_file_key, local_credentials_path)
        logger.info(f"Credentials file downloaded to {local_credentials_path}")
    except boto3.exceptions.S3UploadFailedError as e:
        logger.error(f"Error downloading credentials file: {e}")
        sys.exit(1)

# Download credentials
download_credentials()

# Load external account credentials
credentials, project = load_credentials_from_file(local_credentials_path)

# Specify the target service account to impersonate
target_principal = 'myapp@-nonprod.iam.gserviceaccount.com'
target_scopes = ['https://www.googleapis.com/auth/cloud-platform']

# Create impersonated credentials
impersonated_creds = impersonated_credentials.Credentials(
    source_credentials=credentials,
    target_principal=target_principal,
    target_scopes=target_scopes
)

# Refresh the impersonated credentials
request = Request()
impersonated_creds.refresh(request)

# Initialize the BigQuery client
client = bigquery.Client(credentials=impersonated_creds, project=project)

# Define your DynamicFrames
def create_dynamic_frame(table_name):
    return glueContext.create_dynamic_frame.from_options(
        connection_type="mysql",
        connection_options={
            "useConnectionProperties": "true",
            "dbtable": table_name,
            "connectionName": "myapp nonprod DB",
        },
        transformation_ctx=f"{table_name}_dynamic_frame"
    )

tables = ['changes', 'pull_requests', 'event_change', 'team_members']
dynamic_frames = {table: create_dynamic_frame(table) for table in tables}

# Drop duplicates
def drop_duplicates(frame):
    return DynamicFrame.fromDF(frame.toDF().dropDuplicates(), glueContext, f"{frame.transformation_ctx}_drop_duplicates")

dynamic_frames = {table: drop_duplicates(df) for table, df in dynamic_frames.items()}

# Drop fields for specific tables
drop_fields_frames = {
    'pull_requests': DropFields.apply(frame=dynamic_frames['pull_requests'], paths=["organization_name", "comment_id"], transformation_ctx="pull_requests_drop_fields")
}

# Function to load data to BigQuery
def load_to_bigquery(df, table_id):
    pandas_df = df.toDF().toPandas()
    job_config = bigquery.LoadJobConfig(write_disposition=bigquery.WriteDisposition.WRITE_APPEND)
    load_job = client.load_table_from_dataframe(pandas_df, table_id, job_config=job_config)
    load_job.result()  # Wait for the job to complete
    logger.info(f"Loaded {load_job.output_rows} rows into {table_id}.")

# Write data to BigQuery
for table, df in dynamic_frames.items():
    table_id = f'change_tracker.{table}'
    load_to_bigquery(df, table_id)

for table, df in drop_fields_frames.items():
    table_id = f'change_tracker.{table}'
    load_to_bigquery(df, table_id)

job.commit()

2. Google工作负载身份联合凭证配置

{
  "type": "external_account",
  "audience": "//iam.googleapis.com/projects/123455678/locations/global/workloadIdentityPools/mynonprodpool/providers/aws-devendi-nonprod",
  "subject_token_type": "urn:ietf:params:aws:token-type:aws4_request",
  "service_account_impersonation_url": "https://iamcredentials.googleapis.com/v1/projects/-/serviceAccounts/myapp@nonprod.iam.gserviceaccount.com:generateAccessToken",
  "token_url": "https://sts.googleapis.com/v1/token",
  "credential_source": {
    "environment_id": "aws1",
    "region_url": "http://169.254.169.254/latest/meta-data/placement/availability-zone",
    "url": "http://169.254.169.254/latest/meta-data/iam/security-credentials",
    "regional_cred_verification_url": "https://sts.{region}.amazonaws.com?Action=GetCallerIdentity&Version=2011-06-15"
  }
}

3. 错误信息

Error Category: UNCLASSIFIED_ERROR; RefreshError: ('Unable to acquire impersonated credentials', '{
 "error": {
 "code": 403,
 "message": "Permission \'iam.serviceAccounts.getAccessToken\' denied on resource (or it may not exist).",
 "status": "PERMISSION_DENIED",
 "details": [
 {
 "@type": "type.googleapis.com/google.rpc.ErrorInfo",
 "reason": "IAM_PERMISSION_DENIED",
 "domain": "iam.googleapis.com",
 "metadata": {
 "permission": "iam.serviceAccounts.getAccessToken"
 }
 }
 ]
 }
}\n')

4. AWS IAM角色信任策略

{
"Version": "2012-10-17",
"Statement": [
    {
        "Effect": "Allow",
        "Principal": {
            "Service": "glue.amazonaws.com"
        },
        "Action": "sts:AssumeRole"
    },
    {
        "Effect": "Allow",
        "Principal": {
            "Federated": "arn:aws:iam::123123123:oidc-provider/oidc.googleapis.com"
        },
        "Action": "sts:AssumeRoleWithWebIdentity",
        "Condition": {
            "StringEquals": {
                "oidc.googleapis.com:aud": "aws-devendi-nonprod"
            }
        }
    }
]
}

问题排查与解决方案

核心原因

通过工作负载身份联合认证后的身份,没有被授予impersonate目标Google服务账号的权限,导致无法调用iam.serviceAccounts.getAccessToken接口生成凭证。

修复步骤

  1. 给工作负载身份池中的AWS身份授予服务账号 impersonation 权限

    • 登录Google Cloud Console,找到目标服务账号myapp@nonprod.iam.gserviceaccount.com
    • 进入服务账号的权限页面,点击添加权限
    • 在主体中填入:principal://iam.googleapis.com/projects/[GCP_PROJECT_ID]/locations/global/workloadIdentityPools/mynonprodpool/subject/[AWS_GLUE_ROLE_ARN]
      • 替换[GCP_PROJECT_ID]为你的GCP项目ID(如123455678)
      • 替换[AWS_GLUE_ROLE_ARN]为Glue Job使用的IAM角色完整ARN(比如arn:aws:iam::123123123:role/your-glue-job-execution-role)
    • 在角色中搜索并选择Service Account Token Creator角色(该角色包含所需的iam.serviceAccounts.getAccessToken权限)
    • 保存配置
  2. 修正脚本中的服务账号邮箱错误
    脚本中target_principal字段多了一个多余的减号,需从myapp@-nonprod.iam.gserviceaccount.com修正为myapp@nonprod.iam.gserviceaccount.com,与凭证配置中的服务账号保持一致。

  3. 验证AWS IAM角色的S3访问权限
    确保Glue Job使用的IAM角色拥有s3:GetObject权限,允许访问S3桶aws-glue-assets-123123123-eu-central-1下的credentials/service-account-impersonation.json文件。


内容的提问来源于stack exchange,提问作者Marko Vasic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 12:35:55