You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Glue写入Iceberg表触发Spark内存及Executor失败问题求助

问题:S3读取Parquet创建Iceberg表时任务失败

我从S3落地桶读取总计80个Parquet文件,尝试在原始桶中创建Iceberg表。可以成功生成Dataframe,但执行写入操作时出现以下错误:
1.

ERROR:root:An error occurred while calling o748.defaultParallelism. : java.lang.IllegalStateException: Cannot call methods on a stopped SparkContext. This stopped SparkContext was created at:
Write failed with error: An error occurred while calling o743.create:org.apache.spark.SparkException: Writing job aborted
Caused by: org.apache.spark.SparkException: Job aborted due to stage failure: Task 31 in stage 43.0 failed 4 times, most recent failure: Lost task 31.3 in stage 43.0 (TID 528) (10.236.38.189 executor 19): ExecutorLostFailure (executor 19 exited caused by one of the running tasks) Reason: Remote RPC client disassociated. Likely due to containers exceeding thresholds, or network issues. Check driver logs for WARN messages. Driver stacktrace:

已尝试以下操作但问题未解决:

  • 增加Worker数量
  • 对Dataframe执行重分区操作
  • 执行Dataframe的count()方法时,同样触发上述错误
  • 同一脚本可正常处理其他200张更大的表,仅该表失败

相关代码片段

读取数据代码:

s3_src_landing_ddf = self.glue_context.create_dynamic_frame.from_options(
            format_options={},
            connection_type="s3",
            format="parquet",
            connection_options={
                "paths": [s3_file_list],
                "recurse": recurse_val
            },
            transformation_ctx="s3_src_landing_ddf",
        )

写入Iceberg表代码:

s3_src_landing_df.writeTo("glue_catalog."+schema_nm+"." + table.lower()).tableProperty(
                    "format-version", "2"
                ).tableProperty(
                    "location", "s3://raw_bucket/source_system/db_name+/table.lower()"
                ).create()

触发错误的count()方法代码:

s3_src_landing_df.count()

内容的提问来源于stack exchange,提问作者KarthiK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 00:47:25