You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Glue向Aurora Postgres插入UUID失败问题求助

问题分析与解决方案

问题原因

这不是AWS Glue的Bug,而是数据类型映射规则和Spark序列化机制导致的:

  1. 当传入str(uuid.uuid4())时,Glue默认将该列识别为String类型,JDBC驱动会按character varying类型发送给Postgres。Postgres的严格类型检查会拒绝隐式转换,而SQL客户端能成功是因为SQL解析层会自动处理字符串到UUID的隐式转换,但JDBC驱动默认不开启这个行为。
  2. 直接传入uuid.uuid4()生成的Python UUID对象时,Spark/Glue的序列化器无法正确处理该对象的内部结构,触发递归调用栈溢出,导致RecursionError。

解决方案

方案1:添加JDBC驱动参数开启隐式转换

修改Postgres连接参数,添加stringtype=unspecified,让JDBC驱动将字符串视为未指定类型,触发Postgres的隐式转换:

import uuid
from awsglue.context import GlueContext
from pyspark.context import SparkContext

sc = SparkContext.getOrCreate()
glueContext = GlueContext(sc)

schema = ['id']
rdd = [[str(uuid.uuid4())]]
dyf = glueContext.create_dynamic_frame_from_rdd(rdd, 'dyf', schema=schema)

connection_options = {
    "url": "jdbc:postgresql://your-aurora-endpoint:5432/your-db",
    "dbtable": "data.t",
    "user": "your-username",
    "password": "your-password",
    "stringtype": "unspecified"  # 关键参数
}

glueContext.write_from_options(
    frame_or_dfc=dyf,
    connection_type='postgresql',
    connection_options=connection_options
)

方案2:显式转换数据类型(Spark DataFrame方式)

将Dynamic Frame转换为Spark DataFrame,显式把字符串列转换为UUID类型后写入:

import uuid
from awsglue.context import GlueContext
from pyspark.context import SparkContext

sc = SparkContext.getOrCreate()
glueContext = GlueContext(sc)
spark = glueContext.spark_session

# 创建带明确schema的DataFrame
from pyspark.sql.types import StructType, StructField, StringType
schema = StructType([StructField("id", StringType(), nullable=True)])
rdd = [[str(uuid.uuid4())]]
df = spark.createDataFrame(rdd, schema=schema)

# 转换为UUID类型(Spark 3.x及以上支持)
df = df.withColumn("id", df["id"].cast("uuid"))

# 写入Postgres
df.write \
    .format("jdbc") \
    .option("url", "jdbc:postgresql://your-aurora-endpoint:5432/your-db") \
    .option("dbtable", "data.t") \
    .option("user", "your-username") \
    .option("password", "your-password") \
    .mode("append") \
    .save()

方案3:使用Glue ApplyMapping映射类型

通过ApplyMapping将String类型映射到Postgres的UUID类型:

import uuid
from awsglue.context import GlueContext
from pyspark.context import SparkContext
from awsglue.transforms import ApplyMapping
from pyspark.sql.types import StructType, StructField, StringType

sc = SparkContext.getOrCreate()
glueContext = GlueContext(sc)
spark = glueContext.spark_session

# 定义明确schema
schema = StructType([StructField("id", StringType(), nullable=True)])
rdd = [[str(uuid.uuid4())]]
df = spark.createDataFrame(rdd, schema=schema)
dyf = glueContext.create_dynamic_frame_from_df(df, "dyf")

# 映射String到UUID类型
mapped_dyf = ApplyMapping.apply(
    frame=dyf,
    mappings=[("id", "string", "id", "uuid")]
)

# 写入Postgres
connection_options = {
    "url": "jdbc:postgresql://your-aurora-endpoint:5432/your-db",
    "dbtable": "data.t",
    "user": "your-username",
    "password": "your-password"
}

glueContext.write_from_options(
    frame_or_dfc=mapped_dyf,
    connection_type='postgresql',
    connection_options=connection_options
)

内容的提问来源于stack exchange,提问作者Ken Zhang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 00:10:49