能否通过JDBC访问作为元数据存储的AWS Glue表?
Hey there! Great question—let’s clear this up right away: AWS Glue doesn’t provide a native, direct JDBC endpoint like Hive does for connecting to its Data Catalog. But don’t worry, there are two reliable workarounds to achieve exactly what you need: querying Glue tables via JDBC.
Option 1: Use Amazon Athena with its JDBC Driver
This is the most straightforward, managed approach. Athena natively integrates with the AWS Glue Data Catalog, meaning it uses Glue’s table definitions to query the underlying data (usually stored in S3). Here’s how to set it up:
- First, make sure your Glue tables are properly configured (correct schema, data location in S3, and format like Parquet/CSV).
- Download the official Athena JDBC driver from AWS’s documentation.
- Configure your JDBC client tool (like DBeaver, Tableau, or custom Java apps) with the driver. You’ll need to provide:
- AWS credentials (IAM access key/secret, or use IAM roles if running on AWS services)
- Your AWS region
- An S3 bucket path to store Athena query results (required for Athena to work)
- Once connected, you can run standard SQL queries against your Glue tables just like you would with any JDBC data source.
Option 2: Use Spark Thrift Server with Glue Data Catalog as Metastore
If you need more control over the query engine (like using Spark-specific features), you can set up a Spark Thrift Server that uses Glue as its metastore:
- You can either use an AWS Glue Development Endpoint (a managed Spark environment) or spin up a self-managed Spark cluster (like on EC2 or EMR).
- Configure Spark to use the Glue Data Catalog as its Hive metastore by setting the appropriate Spark configurations, e.g.,
spark.hadoop.hive.metastore.client.factory.classtocom.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory. - Enable the Spark Thrift Server on your cluster/endpoint.
- Use the standard Hive JDBC driver to connect to the Thrift Server, and you’ll be able to query your Glue tables through Spark.
Key Notes
- The Athena method is ideal if you want a serverless, low-maintenance solution—no servers to manage, just pay for the data scanned by your queries.
- The Spark Thrift Server approach gives you more flexibility (like custom Spark UDFs or complex transformations) but requires you to maintain the underlying cluster/endpoint.
内容的提问来源于stack exchange,提问作者Rahul Sharma

