本地PySpark无法访问AWS Data Catalog中非Iceberg表的问题求助
Root Cause
Your current SparkSession config only sets up an Iceberg-specific catalog (glue_catalog) for AWS Glue, but doesn't configure Spark's default Hive catalog to connect to Glue. Non-Iceberg tables reside in Glue's Data Catalog, but Spark is trying to look them up in a local Hive metastore (which lacks these tables), resulting in the TABLE_OR_VIEW_NOT_FOUND error.
Fix Steps
1. Add Required Dependencies
Include the AWS Glue Data Catalog Hive Client in your spark.jars.packages to enable Spark's Hive catalog to communicate with Glue:
Add com.amazonaws:aws-glue-datacatalog-hive-client:1.11.300 to the package list (compatible with your existing AWS SDK bundle version).
2. Configure Hive Catalog to Use Glue
Add these configs to your SparkSession builder to point the default Hive catalog to AWS Glue:
spark.hadoop.hive.metastore.client.factory.class: Specifies the Glue Hive client factoryspark.hadoop.hive.metastore.region: AWS region where your Glue Data Catalog is hosted- (Optional)
spark.hadoop.hive.metastore.glue.catalogid: Your AWS account ID if working across multiple accounts
3. Modified SparkSession Config
Here's the updated code with all necessary changes:
spark = SparkSession.builder.config("spark.hadoop.fs.s3a.access.key", os.getenv('AWS_ACCESS_KEY_ID')) \ .config("spark.hadoop.fs.s3a.secret.key", os.getenv('AWS_SECRET_ACCESS_KEY')) \ .config("spark.hadoop.fs.s3a.session.token", os.getenv('AWS_SESSION_TOKEN')) \ .config("spark.hadoop.fs.s3a.endpoint", f"s3.{os.getenv('AWS_REGION')}.amazonaws.com") \ .config("spark.hadoop.fs.s3a.region", os.getenv('AWS_REGION')) \ .config("spark.jars.packages", "org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.6.0,org.apache.hadoop:hadoop-aws:3.3.4,com.amazonaws:aws-java-sdk-bundle:1.11.874,software.amazon.awssdk:bundle:2.27.2,software.amazon.awssdk:metrics-spi:2.27.2,software.amazon.awssdk:s3:2.27.2,software.amazon.awssdk:glue:2.27.2,com.amazonaws:aws-glue-datacatalog-hive-client:1.11.300") \ .config('spark.sql.catalog.glue_catalog', 'org.apache.iceberg.spark.SparkCatalog') \ .config('spark.sql.catalog.glue_catalog.catalog-impl', 'org.apache.iceberg.aws.glue.GlueCatalog') \ .config('spark.sql.iceberg.handle-timestamp-without-timezone', 'true') \ .config('spark.sql.catalog.glue_catalog.warehouse', 's3://glue/datalake/') \ .config('spark.sql.extensions','org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions')\ .config('spark.sql.catalog.glue_catalog.io-impl', 'org.apache.iceberg.aws.s3.S3FileIO') \ .config("spark.hadoop.fs.s3a.impl", "org.apache.hadoop.fs.s3a.S3AFileSystem") \ .config("spark.hadoop.fs.defaultFS", "s3://glue") \ # Glue Hive metastore configs .config("spark.hadoop.hive.metastore.client.factory.class", "com.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory") \ .config("spark.hadoop.hive.metastore.region", os.getenv('AWS_REGION')) \ # Optional: Uncomment for multi-account setups # .config("spark.hadoop.hive.metastore.glue.catalogid", os.getenv('AWS_ACCOUNT_ID')) \ .enableHiveSupport() \ .getOrCreate()
4. Querying Tables Correctly
- Iceberg tables: Explicitly use the
glue_catalog:SELECT * FROM glue_catalog.your_database.your_iceberg_table; - Non-Iceberg tables: Use the default Hive catalog (no explicit catalog required, or reference
hive_catalog):SELECT * FROM your_database.your_non_iceberg_table; -- OR SELECT * FROM hive_catalog.your_database.your_non_iceberg_table;
Verification
Run these commands to confirm both catalogs are accessible:
# List Iceberg tables in glue_catalog spark.sql("SHOW TABLES IN glue_catalog.your_database").show() # List non-Iceberg tables in default Hive catalog spark.sql("SHOW TABLES IN your_database").show()
内容的提问来源于stack exchange,提问作者Sayantan Banerjee

