You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何查询AWS Glue(含Oracle连接场景)中使用的PySpark版本?

How to Check PySpark Version in Two AWS Glue Scenarios

Great question! Let's walk through how to check your PySpark version in both standard AWS Glue environments and when you're using Glue + PySpark to connect to an Oracle database. Here are the straightforward methods for each scenario:

1. Standard AWS Glue Environment

You have two easy ways to get the PySpark version here:

Option 1: Print the version directly in your Glue job script

Just add one of these lines to your PySpark code, run the job, and check the logs for the output:

# Using the pre-initialized Spark session in Glue jobs
print(f"PySpark Version: {spark.version}")

# Or importing pyspark directly
import pyspark
print(f"PySpark Version: {pyspark.__version__}")

Option 2: Check via the AWS Glue Console (no job run needed)

If you want to confirm the version tied to your job's runtime without executing the script:

  • Go to the AWS Glue Console → Jobs
  • Select your target job and click Edit
  • Navigate to the Job details tab
  • Look for the Glue version field (e.g., Glue 4.0, Glue 3.0)
  • Map this to the corresponding PySpark version (as of 2024):
    • Glue 4.0 → PySpark 3.3.x
    • Glue 3.0 → PySpark 3.1.x
    • Glue 2.0 → PySpark 2.4.x

2. AWS Glue + PySpark Connecting to Oracle Database

The good news is that connecting to Oracle doesn't change how you check the PySpark version—your job's underlying PySpark version is still determined by the Glue runtime you're using. You can use the same methods as above, plus verify it while working with your Oracle connection:

Check in your Oracle-connected job script

Even when reading/writing to an Oracle database, you can print the version at the start of your script to confirm:

# The Spark session is pre-initialized in Glue jobs, but if you need to explicitly create it:
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("OracleGlueJob").getOrCreate()

# Check PySpark version first
print(f"PySpark Version for Oracle Job: {spark.version}")

# Proceed with your Oracle connection logic
oracle_df = spark.read.format("jdbc") \
    .option("url", "jdbc:oracle:thin:@//your-oracle-endpoint:1521/your-db-name") \
    .option("dbtable", "your_schema.your_table") \
    .option("user", "your-username") \
    .option("password", "your-password") \
    .option("driver", "oracle.jdbc.driver.OracleDriver") \
    .load()

# Continue with your data transformations...

Verify via the Glue Console

Same as the standard scenario: head to your job's Job details tab, check the Glue version, and map it to the corresponding PySpark version. The Oracle connection doesn't alter this runtime version.

内容的提问来源于stack exchange,提问作者Bala

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:51:54