AWS Glue从CSV生成结构化JSON出现转义斜杠与字符串化问题求解
问题根因
- 代码中提前使用
to_json将每行数据转换为JSON字符串,collect_list最终收集到的是字符串数组,而非结构化的对象数组 - 后续将数组强制转换为string类型,写入JSON时Spark会对字符串内的特殊字符、双引号做二次转义,最终生成多余转义斜杠,且数组内对象被识别为字符串。
解决方案
直接收集结构化结构体数组,交由Spark JSON写入接口自动序列化,无需提前手动转JSON字符串,修正后代码如下:
from pyspark.sql.functions import struct, collect_list, spark_partition_id SourceDataDYF = glueContext.create_dynamic_frame.from_options( format_options = {"quoteChar": '"', "escaper":"", "withHeader":True, "separator":"|", "inferSchema":"false"}, connection_type = "s3", format = "csv", connection_options = {"paths": ["s3://bucket_name/csv_file_path/"], "recurse":True}, transformation_ctx = "SourceDataDYF" ) StageDataDF = SourceDataDYF.toDF() PreStageDataDF1 = StageDataDF.groupBy(spark_partition_id()) \ .agg(collect_list(struct(*StageDataDF.columns)).alias("awsservices")) \ .select("awsservices").coalesce(1) targetDataDYF = DynamicFrame.fromDF(PreStageDataDF1,glueContext,"PreStageDataDF1") targetDataJSON = glueContext.write_dynamic_frame.from_options( frame = targetDataDYF, connection_type = "s3", connection_options = {"path": "s3://result_bucket_name/folder_path/", "partitionKeys": []}, format = "json", transformation_ctx = "targetDataJSON" )
补充说明
如果CSV读取到的字段类型需要调整(比如需要将字符串转数值类型),可以在collect_list之前先对指定字段做类型转换,再打包进struct结构体即可,不需要额外处理转义逻辑。
内容的提问来源于stack exchange,提问作者Harish Wable
相关产品推荐
相关产品推荐

