You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Glue从CSV生成结构化JSON出现转义斜杠与字符串化问题求解

问题根因
  • 代码中提前使用to_json将每行数据转换为JSON字符串,collect_list最终收集到的是字符串数组,而非结构化的对象数组
  • 后续将数组强制转换为string类型,写入JSON时Spark会对字符串内的特殊字符、双引号做二次转义,最终生成多余转义斜杠,且数组内对象被识别为字符串。
解决方案

直接收集结构化结构体数组,交由Spark JSON写入接口自动序列化,无需提前手动转JSON字符串,修正后代码如下:

from pyspark.sql.functions import struct, collect_list, spark_partition_id

SourceDataDYF = glueContext.create_dynamic_frame.from_options(
   format_options = {"quoteChar": '"', "escaper":"", "withHeader":True, "separator":"|", "inferSchema":"false"},
   connection_type = "s3",
   format = "csv",
   connection_options = {"paths": ["s3://bucket_name/csv_file_path/"], "recurse":True},
   transformation_ctx = "SourceDataDYF"
)

StageDataDF = SourceDataDYF.toDF()

PreStageDataDF1 = StageDataDF.groupBy(spark_partition_id()) \
   .agg(collect_list(struct(*StageDataDF.columns)).alias("awsservices")) \
   .select("awsservices").coalesce(1)

targetDataDYF = DynamicFrame.fromDF(PreStageDataDF1,glueContext,"PreStageDataDF1")
targetDataJSON = glueContext.write_dynamic_frame.from_options(
   frame = targetDataDYF,
   connection_type = "s3",
   connection_options = {"path": "s3://result_bucket_name/folder_path/", "partitionKeys": []},
   format = "json",
   transformation_ctx = "targetDataJSON"
)
补充说明

如果CSV读取到的字段类型需要调整(比如需要将字符串转数值类型),可以在collect_list之前先对指定字段做类型转换,再打包进struct结构体即可,不需要额外处理转义逻辑。

内容的提问来源于stack exchange,提问作者Harish Wable

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 10:36:01