You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Glue作业遇KeyError:如何解决列大小写敏感映射问题?

Fixing AWS Glue KeyError: 'userid' Case Sensitivity Issue

Hey there, let's work through this case sensitivity problem you're hitting in AWS Glue. The error KeyError: 'userid' does not exist tells us that Glue can't find the exact case-matching column name in your dynamic frame, even after setting up the mapping. Let's break down the fixes step by step:

1. Verify the Exact Column Names from Your Dynamic Frame

Before adjusting your mapping, you need to confirm what column names Glue is actually picking up from the CSV (via the crawler). Sometimes crawlers might infer column names with unexpected casing, or there might be typos in your mapping.

Add this line right after you create dynamic_frame_read to print the schema:

print(dynamic_frame_read.printSchema())

Run the job once (or use a Glue interactive session to test this faster) to see the exact case of each column name in the raw dynamic frame. For example, if the raw column is UserId instead of userid, your mapping source needs to match that exactly.

2. Fix Your Mapping Schema and Case Sensitivity Setting

Looking at your current mappingsSchema, there's a potential inconsistent casing in one entry:

('preferredfullName', 'string', 'preferredFullname', 'string')

Make sure every source column name in your mapping exactly matches what's in the raw dynamic frame (case included).

If your goal is to map between different casings without worrying about exact matches, set case_sensitive=False in apply_mapping—this tells Glue to ignore case when matching source columns:

mapped_dynamic_frame_read = dynamic_frame_read.apply_mapping(
    mappings = mappingsSchema,
    case_sensitive = False,
    transformation_ctx = "tfx"
)

3. Bypass the Crawler Schema (Manual CSV Schema Definition)

If the crawler's inferred schema is causing persistent issues, skip relying on it and define the CSV schema manually when reading the file. This gives you full control over column names and types:

from awsglue.dynamicframe import DynamicFrame
from pyspark.sql.types import StructType, StructField, StringType, IntegerType

# Define your target camelCase schema
custom_schema = StructType([
    StructField("userId", IntegerType(), True),
    StructField("jobTitleName", StringType(), True),
    StructField("firstName", StringType(), True),
    StructField("lastName", StringType(), True),
    StructField("preferredFullName", StringType(), True),
    StructField("employeeCode", StringType(), True),
    StructField("region", StringType(), True)
])

# Read CSV directly with your custom schema
df = spark.read.csv(
    "s3://your-bucket-path/your-file.csv",
    header=True,  # Enable this if your CSV has a header row
    schema=custom_schema,
    inferSchema=False  # Disable auto-inference to use your custom schema
)

# Convert back to DynamicFrame if you need Glue-specific transformations
dynamic_frame = DynamicFrame.fromDF(df, glueContext, "custom_dynamic_frame")

This approach avoids any crawler-related casing inconsistencies entirely.

4. Double-Check Your CSV Header Row

Don't overlook the basics: confirm the actual CSV file's header row uses the exact casing you're expecting. If the CSV header has UserId instead of userid, the crawler will infer that as the column name, leading to the KeyError when you reference userid in your mapping.


内容的提问来源于stack exchange,提问作者umdev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 12:07:46