如何查询AWS Glue(含Oracle连接场景)中使用的PySpark版本?
Great question! Let's walk through how to check your PySpark version in both standard AWS Glue environments and when you're using Glue + PySpark to connect to an Oracle database. Here are the straightforward methods for each scenario:
1. Standard AWS Glue Environment
You have two easy ways to get the PySpark version here:
Option 1: Print the version directly in your Glue job script
Just add one of these lines to your PySpark code, run the job, and check the logs for the output:
# Using the pre-initialized Spark session in Glue jobs print(f"PySpark Version: {spark.version}") # Or importing pyspark directly import pyspark print(f"PySpark Version: {pyspark.__version__}")
Option 2: Check via the AWS Glue Console (no job run needed)
If you want to confirm the version tied to your job's runtime without executing the script:
- Go to the AWS Glue Console → Jobs
- Select your target job and click Edit
- Navigate to the Job details tab
- Look for the Glue version field (e.g., Glue 4.0, Glue 3.0)
- Map this to the corresponding PySpark version (as of 2024):
- Glue 4.0 → PySpark 3.3.x
- Glue 3.0 → PySpark 3.1.x
- Glue 2.0 → PySpark 2.4.x
2. AWS Glue + PySpark Connecting to Oracle Database
The good news is that connecting to Oracle doesn't change how you check the PySpark version—your job's underlying PySpark version is still determined by the Glue runtime you're using. You can use the same methods as above, plus verify it while working with your Oracle connection:
Check in your Oracle-connected job script
Even when reading/writing to an Oracle database, you can print the version at the start of your script to confirm:
# The Spark session is pre-initialized in Glue jobs, but if you need to explicitly create it: from pyspark.sql import SparkSession spark = SparkSession.builder.appName("OracleGlueJob").getOrCreate() # Check PySpark version first print(f"PySpark Version for Oracle Job: {spark.version}") # Proceed with your Oracle connection logic oracle_df = spark.read.format("jdbc") \ .option("url", "jdbc:oracle:thin:@//your-oracle-endpoint:1521/your-db-name") \ .option("dbtable", "your_schema.your_table") \ .option("user", "your-username") \ .option("password", "your-password") \ .option("driver", "oracle.jdbc.driver.OracleDriver") \ .load() # Continue with your data transformations...
Verify via the Glue Console
Same as the standard scenario: head to your job's Job details tab, check the Glue version, and map it to the corresponding PySpark version. The Oracle connection doesn't alter this runtime version.
内容的提问来源于stack exchange,提问作者Bala

