调用o60.getDynamicFrame出错:AWS Glue读取Redshift表问题求助
Hey there, let's dig into this Py4JJavaError you're facing when trying to build a DataFrame from your Redshift table. This error usually stems from mismatched connection details, missing dependencies, or permission gaps between your execution environment (like Glue/EMR) and Redshift. Here's how to troubleshoot it step by step:
Double-check your Redshift connection details
It sounds basic, but typos here are super common. Verify:- Your table and schema names are correct—Redshift is case-sensitive unless you wrapped them in double quotes when creating the table.
- Your JDBC URL follows the right format:
jdbc:redshift://<cluster-endpoint>:<port>/<db-name>?user=<username>&password=<password> - Your Redshift cluster's security group allows incoming traffic from your Spark/Glue environment on port 5439 (the default Redshift port).
Ensure the Redshift JDBC driver is properly loaded
Without the right driver, your code can't talk to Redshift. Depending on your setup:- For Glue jobs: Add
--extra-jars s3://your-bucket/path/to/RedshiftJDBC42.jarto your job parameters to pull the driver from S3. - For EMR/self-managed Spark: Make sure the driver JAR is in Spark's
jarsdirectory, or pass it via the--jarsflag when starting your Spark session.
- For Glue jobs: Add
Validate your DynamicFrame logic
If you're using Glue'sgetDynamicFramemethod:- Confirm the database and table names in
from_catalogmatch exactly what's in the Glue Data Catalog (and that the catalog is synced with Redshift). - If you're using direct JDBC instead of the catalog, avoid mixing Glue DynamicFrame methods with raw Spark JDBC calls—stick to one approach to prevent conflicts.
- Confirm the database and table names in
Get the full Java error stack trace
The log snippet you shared cuts off atjava.lang...—the rest of that error message is the real clue. Look for:java.sql.SQLException: Could mean bad credentials, missing table permissions, or a connection timeout.ClassNotFoundException: Definitely means the JDBC driver isn't being loaded correctly.AmazonServiceException: Points to IAM permission issues (like your execution role not having access to Redshift or S3).
Check IAM and Redshift permissions
Make sure the IAM role running your script (Glue job role, EMR instance role) has:- Permissions to describe Redshift clusters (
redshift:DescribeClusters) if using the Glue Catalog. - Access to S3 if your JDBC driver is stored there (
s3:GetObject). - And don't forget: The Redshift database user you're connecting with needs
SELECTpermissions on the target table.
- Permissions to describe Redshift clusters (
Here's a quick example of a working Glue job that pulls from Redshift and converts to a DataFrame:
import sys from awsglue.transforms import * from awsglue.utils import getResolvedOptions from pyspark.context import SparkContext from awsglue.context import GlueContext from awsglue.job import Job # Initialize job context args = getResolvedOptions(sys.argv, ['JOB_NAME']) sc = SparkContext() glueContext = GlueContext(sc) spark = glueContext.spark_session job = Job(glueContext) job.init(args['JOB_NAME'], args) # Pull data from Redshift via Glue Catalog redshift_dynamic_frame = glueContext.create_dynamic_frame.from_catalog( database="your_redshift_catalog_db", table_name="your_target_table", transformation_ctx="redshift_dynamic_frame" ) # Convert to Spark DataFrame redshift_df = redshift_dynamic_frame.toDF() redshift_df.show() job.commit()
内容的提问来源于stack exchange,提问作者Andres Urrego Angel

