如何在Spark Scala作业中配置S3请求付费(request-payer)参数以访问Parquet文件?
I see you're hitting an access denied error when trying to access a Requester Pays-enabled S3 bucket via Spark, even after setting fs.s3.useRequesterPaysHeader to true. The issue here is that you're using the S3A filesystem (configured via fs.s3a.impl), but the parameter you're using is for the older, deprecated S3 filesystem implementation.
The Fix: Use S3A-Specific Requester Pays Configuration
For the S3A filesystem (the recommended standard for S3 access in Spark/Hadoop), the correct configuration parameter to enable requester pays is fs.s3a.requester.pays, set to true.
Here's your updated SparkSession code with the corrected parameter:
val spark = SparkSession.builder() .config("spark.driver.maxResultSize", "0") .config("fs.s3a.impl", "org.apache.hadoop.fs.s3a.S3AFileSystem") .config("fs.s3a.requester.pays", "true") // This is the correct parameter for S3A .appName("Test") .getOrCreate()
Additional Notes to Ensure It Works:
- Hadoop Version Compatibility: Make sure your
hadoop-awsdependency is version 2.8 or newer—this parameter was introduced in Hadoop 2.8. Older versions won't recognize it, so you'll need to upgrade if you're on an earlier release. - Cross-Account Permissions: Even with requester pays enabled, your Spark job's IAM identity (role/user) still needs explicit permissions from the bucket owner (e.g.,
s3:GetObject,s3:ListBucket). Double-check that these permissions are granted in the bucket's policy. - Validate Configuration: You can confirm the parameter is loaded correctly by printing Spark's active config with
spark.conf.getAll.foreach(println)—look forfs.s3a.requester.paysset totrue.
This change will make Spark send the x-amz-request-payer: requester header with every S3 request, matching the behavior of your working CLI command (aws s3 ls <s3_path> --request-payer requester) and resolving the access denied error.
内容的提问来源于stack exchange,提问作者Irfan Khan

