如何从Glue开发终端运行Glue脚本?
Absolutely, you can run your test.py Glue script directly from the command line of your Glue Development Endpoint (which is essentially an EMR cluster under the hood), and spark-submit is fully supported here—let me walk you through exactly how to do this:
Step 1: Make Your Script Accessible
First, ensure your script is reachable from the terminal session:
- If you’ve already copied
test.pyto the terminal’s local filesystem, you can use its direct path (e.g.,./test.pyif it’s in your current working directory). - If the script lives in an S3 bucket, you can either copy it to the terminal first with the AWS CLI:
Or reference the S3 path directly in youraws s3 cp s3://your-bucket/path/to/test.py ./spark-submitcommand (Spark can natively read scripts from S3 in EMR environments).
Step 2: Execute the Script with spark-submit
Since the Glue Development Endpoint is built on EMR, it uses the standard Spark runtime, so spark-submit works just like it would on any regular EMR cluster. Use this command (tweak parameters based on your needs):
spark-submit \ --master yarn \ --deploy-mode client \ --conf spark.hadoop.hive.metastore.client.factory.class=com.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory \ ./test.py
Here’s what each key parameter does:
--master yarn: EMR clusters rely on YARN as the cluster manager, so this tells Spark to submit the job to the YARN framework.--deploy-mode client: This runs the Spark driver directly on your terminal, so you’ll see real-time output in the command line. If you want the driver to run on a cluster node instead, switch toclustermode (logs will be stored in your EMR cluster’s default log location or the S3 bucket you configured for the endpoint).- The
--confflag: This configures Spark to use the Glue Data Catalog as the Hive metastore—critical if your script interacts with Glue tables or databases.
Important Notes
- Glue Context Support: If your script uses
GlueContext, don’t worry—the Glue Development Endpoint environment is pre-configured to support it, so no extra setup is needed. - Permissions: Double-check that the IAM role attached to your endpoint has the necessary permissions (e.g., S3 read/write, Glue Data Catalog access) to execute all operations in your script.
- Logging: For
clusterdeploy mode, you can find job logs in the EMR console under your cluster’s "Logs" tab, or in the S3 log bucket you specified when creating the endpoint.
This command-line approach is a totally valid alternative to using the Glue console, giving you more flexibility for testing and running scripts directly from your terminal session.
内容的提问来源于stack exchange,提问作者Lijju Mathew

