如何使用AWS Glue对嵌套结构XML数据源进行编目?
Alright, let's break down how to catalog your nested XML data with AWS Glue— I’ve tackled similar nested XML structures before, so here’s a step-by-step guide that should work for your book catalog example:
First, make sure your XML files are stored in an Amazon S3 bucket. For example, drop them in a path like s3://your-glue-data-bucket/book-catalog/. Keep the structure simple (no overly nested folders unless you need them) so the Glue crawler can easily scan the files.
Glue’s default classifiers might not always perfectly parse nested XML, so creating a custom one ensures it recognizes your structure correctly:
- Head to the AWS Glue Console → Classifiers → Add classifier
- Select XML classifier and name it something like
BookCatalogXMLClassifier - Fill in the parameters:
- Row tag:
book(this tells Glue that each<book>element is a single record) - Root path:
/catalog(points to the top-level container holding all your book records)
- Row tag:
- Save the classifier. This will take priority over Glue’s default classifiers when crawling.
Now create a crawler to scan your XML data and build the catalog:
- Go to Glue Console → Crawlers → Add crawler
- Name your crawler (e.g.,
BookCatalogCrawler) and select an IAM role with permissions to read from your S3 bucket and write to Glue Data Catalog. If you don’t have one, create a new role with theAWSGlueServiceRolepolicy attached, plus S3 read access for your bucket. - For Data source, select S3 and enter the path where your XML files are stored.
- Under Add classifiers, select the custom XML classifier you created earlier.
- Choose a target database (create a new one like
book_catalog_dbif you don’t have an existing one) and set a table prefix (optional, but helps organize tables). - Skip any advanced settings unless you need them, then finish creating the crawler.
- Run the crawler— it should take a minute or two to scan your XML and build the table.
Once the crawler finishes:
- Go to Glue Console → Databases → Select
book_catalog_db→ Open the table created by the crawler - Check the schema: You should see fields like
id(from the book’s attribute),title,genre,price,publish_date,description, and a nestedauthorsfield. Theauthorsfield will be an array containing anauthorstruct, which in turn has anamefield. - Use the Data preview feature to confirm that nested data is being parsed correctly (e.g., each book’s author name shows up under the
authors.author.namepath).
If you need to flatten the nested structure (like turning the authors array into individual rows), you can use a Glue ETL job with PySpark:
from awsglue.context import GlueContext from pyspark.context import SparkContext # Initialize contexts sc = SparkContext() glueContext = GlueContext(sc) # Read the cataloged table book_dyf = glueContext.create_dynamic_frame.from_catalog( database="book_catalog_db", table_name="book_catalog" ) # Unnest the authors array and author struct flattened_dyf = book_dyf.unnest_field("authors") flattened_dyf = flattened_dyf.unnest_field("authors.author") # Preview the flattened data flattened_dyf.show() # If you want to write the flattened data back to the catalog or S3 # glueContext.write_dynamic_frame.from_catalog( # frame=flattened_dyf, # database="book_catalog_db", # table_name="flattened_book_catalog" # )
- XML Attributes: Glue automatically parses XML attributes (like
book id="bk101") as top-level fields in your table— no extra configuration needed. - Missing Fields: If some books have missing fields (e.g., no
description), Glue will infer the field type based on existing values, but you might need to manually adjust the schema later if needed. - Permissions: Double-check your IAM role has
s3:GetObjectaccess to your XML bucket and all Glue Data Catalog permissions (likeglue:CreateTable,glue:UpdateTable).
内容的提问来源于stack exchange,提问作者jsleeuw

