You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark小数据集Join时Executor OOM问题(退出码143)

问题:小数据集缓存后执行左外连接触发OOM

场景说明

  • 处理总大小41MB的小数据集,执行左外连接操作
  • Executor内存配置为10GB
  • 不使用内存持久化(cache()/persist())时任务正常运行;使用cache()或persist(StorageLevel.MEMORY_AND_DISK())时任务持续失败,仅persist(StorageLevel.DISK_ONLY())可正常执行

异常日志

# java.lang.OutOfMemoryError: Java heap space
# -XX:OnOutOfMemoryError="kill %p"
#   Executing /bin/sh -c "kill 42447"...
23/06/14 19:38:57 WARN TaskMemoryManager: Failed to allocate a page (64102862 bytes), try again.
23/06/14 19:38:57 INFO TaskMemoryManager: Memory used in task 4
23/06/14 19:38:57 INFO TaskMemoryManager: 64102862 bytes of memory were used by task 4 but are not associated with specific consumers
23/06/14 19:38:57 INFO TaskMemoryManager: 5500051918 bytes of memory are used for execution and 38002226 bytes of memory are used for storage
23/06/14 19:38:57 WARN BlockManager: Block rdd_27_0 could not be removed as it was not found on disk or in memory

容器终止提示

Diagnostics: [2023-05-28 16:24:44.080]Container killed on request. Exit code is 143
[2023-05-28 16:24:44.080]Container exited with a non-zero exit code 143.

相关截图

  • 任务正常运行截图:任务正常运行截图
  • 任务失败截图:任务失败截图
  • 存储相关截图:存储截图

代码片段

Dataset<Row> currentDataSet = hdfsHandler.loadDataSet(sparkSession, "somelocation");
currentDataSet.createOrReplaceTempView(CURRENT_USER_VIEW);
currentDataSet.cache() //<- 此处引发问题
Dataset<Row> previousDataSet = hdfsHandler.loadDataSet(sparkSession, "somelocation2");
previousDataSet.createOrReplaceTempView(PREVIOUS_USER_VIEW);
previousDataSet.cache() //<- 此处引发问题

String currentRunColumn = "c.".concat("userid");
String previousRunColumn = "p.".concat("userid");

Dataset<Row> addedRecordDataSets = sparkSession.sql("SELECT " + currentRunColumn + " FROM " +
        CURRENT_USER_VIEW + " AS c " +
        " LEFT JOIN " + PREVIOUS_USER_VIEW + " AS p " +
        " ON " + currentRunColumn + " == " + previousRunColumn +
        " WHERE " + previousRunColumn + " IS NULL ");

dataSet.coalesce(1).persist(StorageLevel.DISK_ONLY()).foreachPartition(persist());

疑问与分析

  • 日志显示异常发生时执行内存占用约5GB,因无法分配更多内存导致任务终止
  • 无法理解仅41MB的数据集为何会占用5GB执行内存,询问是否有方法估算Spark任务的执行所需内存(已查阅Spark内存管理相关资料)

更新内容

尽管数据集很小,峰值执行内存达48.1GB,但Executor内存仅3GB时任务并未崩溃;添加内存缓存后,即使Executor内存为10GB也会触发OOM,怀疑是Spark 2.4.x版本的缓存驱逐机制存在Bug。

内容的提问来源于stack exchange,提问作者best wishes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 00:22:43