You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

本地PySpark无法访问AWS Data Catalog中非Iceberg表的问题求助

Solution to Access Both Iceberg and Non-Iceberg Tables in AWS Glue Data Catalog

Root Cause

Your current SparkSession config only sets up an Iceberg-specific catalog (glue_catalog) for AWS Glue, but doesn't configure Spark's default Hive catalog to connect to Glue. Non-Iceberg tables reside in Glue's Data Catalog, but Spark is trying to look them up in a local Hive metastore (which lacks these tables), resulting in the TABLE_OR_VIEW_NOT_FOUND error.

Fix Steps

1. Add Required Dependencies

Include the AWS Glue Data Catalog Hive Client in your spark.jars.packages to enable Spark's Hive catalog to communicate with Glue:
Add com.amazonaws:aws-glue-datacatalog-hive-client:1.11.300 to the package list (compatible with your existing AWS SDK bundle version).

2. Configure Hive Catalog to Use Glue

Add these configs to your SparkSession builder to point the default Hive catalog to AWS Glue:

  • spark.hadoop.hive.metastore.client.factory.class: Specifies the Glue Hive client factory
  • spark.hadoop.hive.metastore.region: AWS region where your Glue Data Catalog is hosted
  • (Optional) spark.hadoop.hive.metastore.glue.catalogid: Your AWS account ID if working across multiple accounts

3. Modified SparkSession Config

Here's the updated code with all necessary changes:

spark = SparkSession.builder.config("spark.hadoop.fs.s3a.access.key", os.getenv('AWS_ACCESS_KEY_ID')) \
            .config("spark.hadoop.fs.s3a.secret.key", os.getenv('AWS_SECRET_ACCESS_KEY')) \
            .config("spark.hadoop.fs.s3a.session.token", os.getenv('AWS_SESSION_TOKEN')) \
            .config("spark.hadoop.fs.s3a.endpoint", f"s3.{os.getenv('AWS_REGION')}.amazonaws.com") \
            .config("spark.hadoop.fs.s3a.region", os.getenv('AWS_REGION')) \
            .config("spark.jars.packages", "org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.6.0,org.apache.hadoop:hadoop-aws:3.3.4,com.amazonaws:aws-java-sdk-bundle:1.11.874,software.amazon.awssdk:bundle:2.27.2,software.amazon.awssdk:metrics-spi:2.27.2,software.amazon.awssdk:s3:2.27.2,software.amazon.awssdk:glue:2.27.2,com.amazonaws:aws-glue-datacatalog-hive-client:1.11.300") \
            .config('spark.sql.catalog.glue_catalog', 'org.apache.iceberg.spark.SparkCatalog') \
            .config('spark.sql.catalog.glue_catalog.catalog-impl', 'org.apache.iceberg.aws.glue.GlueCatalog') \
            .config('spark.sql.iceberg.handle-timestamp-without-timezone', 'true') \
            .config('spark.sql.catalog.glue_catalog.warehouse', 's3://glue/datalake/') \
            .config('spark.sql.extensions','org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions')\
            .config('spark.sql.catalog.glue_catalog.io-impl', 'org.apache.iceberg.aws.s3.S3FileIO') \
            .config("spark.hadoop.fs.s3a.impl", "org.apache.hadoop.fs.s3a.S3AFileSystem") \
            .config("spark.hadoop.fs.defaultFS", "s3://glue") \
            # Glue Hive metastore configs
            .config("spark.hadoop.hive.metastore.client.factory.class", "com.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory") \
            .config("spark.hadoop.hive.metastore.region", os.getenv('AWS_REGION')) \
            # Optional: Uncomment for multi-account setups
            # .config("spark.hadoop.hive.metastore.glue.catalogid", os.getenv('AWS_ACCOUNT_ID')) \
            .enableHiveSupport() \
            .getOrCreate()

4. Querying Tables Correctly

  • Iceberg tables: Explicitly use the glue_catalog:
    SELECT * FROM glue_catalog.your_database.your_iceberg_table;
    
  • Non-Iceberg tables: Use the default Hive catalog (no explicit catalog required, or reference hive_catalog):
    SELECT * FROM your_database.your_non_iceberg_table;
    -- OR
    SELECT * FROM hive_catalog.your_database.your_non_iceberg_table;
    

Verification

Run these commands to confirm both catalogs are accessible:

# List Iceberg tables in glue_catalog
spark.sql("SHOW TABLES IN glue_catalog.your_database").show()

# List non-Iceberg tables in default Hive catalog
spark.sql("SHOW TABLES IN your_database").show()

内容的提问来源于stack exchange,提问作者Sayantan Banerjee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 07:58:24