Glue Job写入S3的Parquet文件数据类型异常问题
Glue Job写入Parquet时列类型为object而非string的解决思路
问题背景
使用Glue Job读取包含JSON数据的manifest文件,转换为DataFrame处理后,通过GLUE_CONTEXT.write_dynamic_frame.from_options将数据以Parquet格式写入S3,但目标Parquet文件中指定列的数据类型被映射为object,而非预期的string。已尝试用df.withColumn("my_column", col("my_column").cast("string"))转换类型,但问题仍存在。
解决思路
确保DataFrame转DynamicFrame时保留类型
若将cast后的DataFrame转为DynamicFrame再写入,转换过程中Glue可能重新推断类型导致object类型回归,需显式指定schema完成转换:from pyspark.sql.types import StringType, StructField, StructType # 定义目标schema,指定my_column为string类型 target_schema = StructType([ # 其他列按实际情况定义 StructField("my_column", StringType(), nullable=True) ]) # 用指定schema将DataFrame转为DynamicFrame transformed_with_contracts = DynamicFrame.fromDF( df, glueContext, "transformed_df", schema=target_schema )直接在DynamicFrame层面转换类型
跳过DataFrame的cast操作,直接对DynamicFrame使用resolveChoice方法强制转换列类型:transformed_with_contracts = transformed_with_contracts.resolveChoice( specs=[("my_column", "cast:string")] )从读取阶段固定列类型
读取JSON数据时直接指定schema,避免Glue自动推断出object类型:from pyspark.sql.types import StructType, StringType, StructField json_schema = StructType([ StructField("my_column", StringType(), nullable=True), # 其他列定义 ]) source_df = glueContext.create_dynamic_frame.from_options( connection_type="s3", connection_options={"paths": ["s3://your-source-path/"]}, format="json", format_options={"schema": json_schema.json()} )检查列内是否存在混合类型
JSON数据中若my_column同时存在字符串、数字等多种类型值,会导致Spark/Glue将其推断为object,需先清洗数据统一列内值类型:# 统一转换为字符串,处理空值 df = df.withColumn("my_column", when(col("my_column").isNull(), "").otherwise(col("my_column").cast("string")))
内容的提问来源于stack exchange,提问作者Nakul
相关产品推荐
相关产品推荐

