AWS Glue作业遇KeyError:如何解决列大小写敏感映射问题?
Hey there, let's work through this case sensitivity problem you're hitting in AWS Glue. The error KeyError: 'userid' does not exist tells us that Glue can't find the exact case-matching column name in your dynamic frame, even after setting up the mapping. Let's break down the fixes step by step:
1. Verify the Exact Column Names from Your Dynamic Frame
Before adjusting your mapping, you need to confirm what column names Glue is actually picking up from the CSV (via the crawler). Sometimes crawlers might infer column names with unexpected casing, or there might be typos in your mapping.
Add this line right after you create dynamic_frame_read to print the schema:
print(dynamic_frame_read.printSchema())
Run the job once (or use a Glue interactive session to test this faster) to see the exact case of each column name in the raw dynamic frame. For example, if the raw column is UserId instead of userid, your mapping source needs to match that exactly.
2. Fix Your Mapping Schema and Case Sensitivity Setting
Looking at your current mappingsSchema, there's a potential inconsistent casing in one entry:
('preferredfullName', 'string', 'preferredFullname', 'string')
Make sure every source column name in your mapping exactly matches what's in the raw dynamic frame (case included).
If your goal is to map between different casings without worrying about exact matches, set case_sensitive=False in apply_mapping—this tells Glue to ignore case when matching source columns:
mapped_dynamic_frame_read = dynamic_frame_read.apply_mapping( mappings = mappingsSchema, case_sensitive = False, transformation_ctx = "tfx" )
3. Bypass the Crawler Schema (Manual CSV Schema Definition)
If the crawler's inferred schema is causing persistent issues, skip relying on it and define the CSV schema manually when reading the file. This gives you full control over column names and types:
from awsglue.dynamicframe import DynamicFrame from pyspark.sql.types import StructType, StructField, StringType, IntegerType # Define your target camelCase schema custom_schema = StructType([ StructField("userId", IntegerType(), True), StructField("jobTitleName", StringType(), True), StructField("firstName", StringType(), True), StructField("lastName", StringType(), True), StructField("preferredFullName", StringType(), True), StructField("employeeCode", StringType(), True), StructField("region", StringType(), True) ]) # Read CSV directly with your custom schema df = spark.read.csv( "s3://your-bucket-path/your-file.csv", header=True, # Enable this if your CSV has a header row schema=custom_schema, inferSchema=False # Disable auto-inference to use your custom schema ) # Convert back to DynamicFrame if you need Glue-specific transformations dynamic_frame = DynamicFrame.fromDF(df, glueContext, "custom_dynamic_frame")
This approach avoids any crawler-related casing inconsistencies entirely.
4. Double-Check Your CSV Header Row
Don't overlook the basics: confirm the actual CSV file's header row uses the exact casing you're expecting. If the CSV header has UserId instead of userid, the crawler will infer that as the column name, leading to the KeyError when you reference userid in your mapping.
内容的提问来源于stack exchange,提问作者umdev

