如何修改Glue Crawler自动检测分区生成的列名?
Got it, let's tackle how to rename those auto-generated partition columns from your Glue Crawler. Since your Kinesis Firehose is writing to S3 with yyyy/mm/dd/HH prefixes, the Crawler defaults to columns like year, month, day, hour—but you can customize these names in a few practical ways, depending on your scenario:
1. Configure the Crawler to Use Custom Partition Names (Recommended: Fix at the Source)
This is the cleanest approach because it tells the Crawler exactly what column names to use from the start, avoiding post-crawl edits. Here's how:
- Head to the AWS Glue Console, find your target Crawler, and click Edit.
- Navigate to the Configure crawler settings step, then expand Advanced options.
- Scroll to the Partitioning section, check Custom partition format.
- Enter a partition path template that maps your S3 prefix structure to your desired column names. For example, if you want columns like
log_year,log_month,log_day,log_hour, use:
Note: If your existing S3 paths don't use thes3://your-bucket/log_year=!{year}/log_month=!{month}/log_day=!{day}/log_hour=!{hour}/key=valueformat (just plainyyyy/mm/dd/HH), you can still define custom names by using!{year}without the key prefix—then the Crawler will assign your specified names to each path segment. - Save the Crawler configuration and re-run it. The new table will have your custom partition column names right away.
2. Rename Partition Columns in an Existing Table
If you already have a table generated by the Crawler and want to update its partition column names directly:
Option A: Use the Glue Console
- Go to the Glue Tables page, select your target table, and click Edit table.
- Switch to the Partition keys tab.
- Simply edit the Name field for each partition key (e.g., change
yeartocustom_year) and save your changes. - To ensure the Crawler doesn't overwrite your edits later, go back to your Crawler's settings, expand Advanced options, and set Update the table definition in the data catalog to Only add new columns.
Option B: Use AWS CLI
If you prefer scripted changes, use the aws glue update-table command to modify the partition keys. Example command:
aws glue update-table --database-name your-database-name --table-input '{ "Name": "your-table-name", "PartitionKeys": [ {"Name": "log_year", "Type": "string"}, {"Name": "log_month", "Type": "string"}, {"Name": "log_day", "Type": "string"}, {"Name": "log_hour", "Type": "string"} ] }'
After running this, your table's partition column names will be updated. Re-run the Crawler to sync the partitions with the new names.
3. Use a Glue ETL Job to Rename Columns (For Bulk Transformations)
If you need to reprocess existing data and rename columns in the process, create a Glue ETL Job with a PySpark script like this:
from awsglue.transforms import RenameField from awsglue.context import GlueContext from pyspark.context import SparkContext from awsglue.job import Job sc = SparkContext() glueContext = GlueContext(sc) spark = glueContext.spark_session job = Job(glueContext) # Read data from the original table original_data = glueContext.create_dynamic_frame.from_catalog( database="your-database-name", table_name="your-original-table" ) # Rename each partition column to your desired names renamed_data = RenameField.apply(frame=original_data, old_name="year", new_name="log_year") renamed_data = RenameField.apply(frame=renamed_data, old_name="month", new_name="log_month") renamed_data = RenameField.apply(frame=renamed_data, old_name="day", new_name="log_day") renamed_data = RenameField.apply(frame=renamed_data, old_name="hour", new_name="log_hour") # Write the transformed data to a new table (or overwrite the original if needed) glueContext.write_dynamic_frame.from_catalog( frame=renamed_data, database="your-database-name", table_name="your-updated-table" ) job.commit()
Run this job, and the output table will have your custom partition column names.
内容的提问来源于stack exchange,提问作者Henrique Barcelos

