Spark小数据集Join时Executor OOM问题(退出码143)
问题:小数据集缓存后执行左外连接触发OOM
场景说明
- 处理总大小41MB的小数据集,执行左外连接操作
- Executor内存配置为10GB
- 不使用内存持久化(
cache()/persist())时任务正常运行;使用cache()或persist(StorageLevel.MEMORY_AND_DISK())时任务持续失败,仅persist(StorageLevel.DISK_ONLY())可正常执行
异常日志
# java.lang.OutOfMemoryError: Java heap space # -XX:OnOutOfMemoryError="kill %p" # Executing /bin/sh -c "kill 42447"... 23/06/14 19:38:57 WARN TaskMemoryManager: Failed to allocate a page (64102862 bytes), try again. 23/06/14 19:38:57 INFO TaskMemoryManager: Memory used in task 4 23/06/14 19:38:57 INFO TaskMemoryManager: 64102862 bytes of memory were used by task 4 but are not associated with specific consumers 23/06/14 19:38:57 INFO TaskMemoryManager: 5500051918 bytes of memory are used for execution and 38002226 bytes of memory are used for storage 23/06/14 19:38:57 WARN BlockManager: Block rdd_27_0 could not be removed as it was not found on disk or in memory
容器终止提示
Diagnostics: [2023-05-28 16:24:44.080]Container killed on request. Exit code is 143 [2023-05-28 16:24:44.080]Container exited with a non-zero exit code 143.
相关截图
- 任务正常运行截图:

- 任务失败截图:

- 存储相关截图:

代码片段
Dataset<Row> currentDataSet = hdfsHandler.loadDataSet(sparkSession, "somelocation"); currentDataSet.createOrReplaceTempView(CURRENT_USER_VIEW); currentDataSet.cache() //<- 此处引发问题 Dataset<Row> previousDataSet = hdfsHandler.loadDataSet(sparkSession, "somelocation2"); previousDataSet.createOrReplaceTempView(PREVIOUS_USER_VIEW); previousDataSet.cache() //<- 此处引发问题 String currentRunColumn = "c.".concat("userid"); String previousRunColumn = "p.".concat("userid"); Dataset<Row> addedRecordDataSets = sparkSession.sql("SELECT " + currentRunColumn + " FROM " + CURRENT_USER_VIEW + " AS c " + " LEFT JOIN " + PREVIOUS_USER_VIEW + " AS p " + " ON " + currentRunColumn + " == " + previousRunColumn + " WHERE " + previousRunColumn + " IS NULL "); dataSet.coalesce(1).persist(StorageLevel.DISK_ONLY()).foreachPartition(persist());
疑问与分析
- 日志显示异常发生时执行内存占用约5GB,因无法分配更多内存导致任务终止
- 无法理解仅41MB的数据集为何会占用5GB执行内存,询问是否有方法估算Spark任务的执行所需内存(已查阅Spark内存管理相关资料)
更新内容
尽管数据集很小,峰值执行内存达48.1GB,但Executor内存仅3GB时任务并未崩溃;添加内存缓存后,即使Executor内存为10GB也会触发OOM,怀疑是Spark 2.4.x版本的缓存驱逐机制存在Bug。
内容的提问来源于stack exchange,提问作者best wishes
相关产品推荐
相关产品推荐

