You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Glue写入Redshift报pyWriteDynamicFrame UNKNOWN数据源错误

问题现象

运行将S3存储的JSON文件上传至Redshift的Glue管道时,作业抛出核心报错:

An error occurred while calling o100.pyWriteDynamicFrame. Failed to find data source: UNKNOWN.

作业日志按顺序输出以下错误信息:

  • InvocationTargetException java.lang.reflect.InvocationTargetException
  • Exception in User Class java.lang.reflect.UndeclaredThrowableException

故障复现使用的Glue作业脚本如下:

import sys
from awsglue.transforms import *
from awsglue.utils import getResolvedOptions
from pyspark.context import SparkContext
from awsglue.context import GlueContext
from awsglue.job import Job

args = getResolvedOptions(sys.argv, ["JOB_NAME"])
sc = SparkContext()
glueContext = GlueContext(sc)
spark = glueContext.spark_session
job = Job(glueContext)
job.init(args["JOB_NAME"], args)

# S3数据源读取节点
S3bucket_node1 = glueContext.create_dynamic_frame.from_options(
    format_options={"multiline": False},
    connection_type="s3",
    format="json",
    connection_options={"paths": ["s3://numbeo-bucket/results.json"], "recurse": True},
    transformation_ctx="S3bucket_node1",
)

# 字段映射节点
ApplyMapping_node2 = ApplyMapping.apply(
    frame=S3bucket_node1, mappings=[], transformation_ctx="ApplyMapping_node2"
)

# Redshift写入节点
RedshiftCluster_node3 = glueContext.write_dynamic_frame.from_catalog(
    frame=ApplyMapping_node2,
    database="redshift-cluster-1",
    table_name="results_json",
    redshift_tmp_dir=args["TempDir"],
    transformation_ctx="RedshiftCluster_node3",
)
故障根因

报错提示找不到UNKNOWN类型的数据源,本质是Glue执行Redshift写入操作时,没有从传入参数、Catalog元数据中识别到合法的Redshift数据源配置,常见触发场景有三类:

  1. 调用write_dynamic_frame.from_catalog方法写入时,Glue Data Catalog中对应的库表没有关联有效的Redshift连接配置,元数据缺失JDBC连接串、集群地址等必填信息,导致Spark无法匹配到对应的Redshift数据源驱动。
  2. 作业使用的Glue版本存在内置连接器缺失问题,或绑定的IAM角色权限不足,无法正常读取Catalog中的连接配置、访问S3临时路径、调用Redshift的COPY命令。
  3. 脚本末尾缺失job.commit()语句,作业执行到收尾阶段无法正常提交写入任务,抛出反射调用异常。
修复步骤
  • 替换Redshift写入节点的实现逻辑,不要依赖Catalog隐式读取连接配置,改为显式指定Redshift连接参数,参考代码如下:
RedshiftCluster_node3 = glueContext.write_dynamic_frame.from_options(
    frame=ApplyMapping_node2,
    connection_type="redshift",
    connection_options={
        "redshiftTmpDir": args["TempDir"],
        "useConnectionProperties": "true",
        "dbtable": "public.results_json", # 替换为Redshift侧实际的schema.表名
        "connectionName": "Glue控制台中提前创建好的Redshift连接名称",
        # 提前配置建表语句,避免表不存在时报错
        "preactions": "CREATE TABLE IF NOT EXISTS public.results_json (根据JSON结构定义对应字段名、字段类型);"
    },
    transformation_ctx="RedshiftCluster_node3",
)
  • 在脚本末尾补充job.commit()语句,保证作业正常提交。
  • 补全ApplyMapping节点的字段映射规则,将S3读取到的JSON字段和Redshift表的字段一一对应,明确字段类型,避免后续写入时出现类型不匹配问题。
  • 检查Glue作业绑定的IAM角色权限,确保角色拥有:S3临时路径的读写权限、Glue Catalog对应库表的访问权限、Redshift集群的连接与数据写入权限。
  • 若坚持使用from_catalog方式写入,需进入Glue Data Catalog控制台,找到redshift-cluster-1库下的results_json表,编辑表属性,关联提前创建好的可用Redshift JDBC连接,补全所有连接必填参数。

内容的提问来源于stack exchange,提问作者Griffin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 17:06:27