Spark2.4+Hadoop2.7读取SSE-KMS加密S3对象失败求助
Hey there, let's work through this issue together. That error message points directly to a signature version conflict, but the real problem is a mix of incompatible library versions and old Hadoop 2.7 jar files overriding your newer dependencies. Here's how to fix it step by step:
1. Fix Library Version Compatibility
Your current aws-java-sdk:1.9.5 is way too old—it doesn’t properly support AWS Signature Version 4 for KMS-encrypted S3 objects, and it’s incompatible with hadoop-aws:3.1.1.
Hadoop 3.1.1 requires a matching AWS SDK version (typically from the 1.11.x series). Update your --packages to use a compatible pair:
os.environ['PYSPARK_SUBMIT_ARGS'] = "--packages=org.apache.hadoop:hadoop-aws:3.1.1,org.apache.hadoop:hadoop-common:3.1.1,org.apache.hadoop:hadoop-auth:3.1.1,com.amazonaws:aws-java-sdk:1.11.760 pyspark-shell"
(Note: 1.11.760 is a tested compatible version for Hadoop 3.1.1; you can cross-check official Hadoop docs for exact version mappings if needed.)
2. Force Classloading Priority to Avoid Old Hadoop Jars
Even with --packages, your system’s Hadoop 2.7 jars might still load first, overriding the newer hadoop-aws classes. Fix this by explicitly setting the classpath to prioritize your new dependencies:
Add these flags to your PYSPARK_SUBMIT_ARGS (or pass them directly in spark-submit):
--conf spark.driver.extraClassPath=$(echo ~/.ivy2/jars/org.apache.hadoop_hadoop-aws-3.1.1.jar:~/.ivy2/jars/com.amazonaws_aws-java-sdk-1.11.760.jar:~/.ivy2/jars/org.apache.hadoop_hadoop-common-3.1.1.jar:~/.ivy2/jars/org.apache.hadoop_hadoop-auth-3.1.1.jar) \ --conf spark.executor.extraClassPath=$(echo ~/.ivy2/jars/org.apache.hadoop_hadoop-aws-3.1.1.jar:~/.ivy2/jars/com.amazonaws_aws-java-sdk-1.11.760.jar:~/.ivy2/jars/org.apache.hadoop_hadoop-common-3.1.1.jar:~/.ivy2/jars/org.apache.hadoop_hadoop-auth-3.1.1.jar) \
(Adjust the paths to match where Ivy downloads your packages—run spark-submit --verbose to find the exact locations.)
3. Clean Up and Correct S3A Configuration
You have redundant or outdated configs cluttering your setup. Update your Spark Hadoop configs to these streamlined settings:
spark._jsc.hadoopConfiguration().set("fs.s3a.access.key", aws_access_id) spark._jsc.hadoopConfiguration().set("fs.s3a.secret.key", aws_access_key) spark._jsc.hadoopConfiguration().set("fs.s3a.fast.upload", "true") spark._jsc.hadoopConfiguration().set("fs.s3a.impl", "org.apache.hadoop.fs.s3a.S3AFileSystem") # Explicitly enforce V4 signing (critical for KMS) spark._jsc.hadoopConfiguration().set("fs.s3a.signing-algorithm", "SIGNING_ALGORITHM_V4") spark._jsc.hadoopConfiguration().set("com.amazonaws.services.s3.enableV4", "true") spark._jsc.hadoopConfiguration().set("fs.s3a.aws.credentials.provider","org.apache.hadoop.fs.s3a.BasicAWSCredentialsProvider") spark._jsc.hadoopConfiguration().set("fs.s3a.endpoint", f"s3.{aws_region}.amazonaws.com") # Simplify SSE-KMS configs spark._jsc.hadoopConfiguration().set("fs.s3a.server-side-encryption-algorithm", aws_sse_algorithm) spark._jsc.hadoopConfiguration().set("fs.s3a.sse.kms.keyId", aws_sse_key)
- Removed redundant
fs.s3a.sse.enabledandfs.s3a.enableServerSideEncryption(the algorithm setting already implies SSE is enabled) - Added
fs.s3a.signing-algorithm—this is required in Hadoop 3.x to explicitly enforce V4 signatures
4. Verify Your Classpath
To confirm the new jars are being loaded instead of the old Hadoop 2.7 ones, add this snippet to your code:
from py4j.java_gateway import java_import java_import(spark._jvm, "java.lang.ClassLoader") class_loader = spark._jvm.Thread.currentThread().getContextClassLoader() hadoop_s3a_class = class_loader.loadClass("org.apache.hadoop.fs.s3a.S3AFileSystem") print(f"Loaded S3AFileSystem from: {hadoop_s3a_class.getProtectionDomain().getCodeSource().getLocation()}")
This should output the path to your hadoop-aws-3.1.1.jar, not the old Hadoop 2.7 version.
Once you apply these changes, your Spark job should properly use V4 signatures and access the KMS-encrypted S3 objects without that error.
内容的提问来源于stack exchange,提问作者mzdrgwj

