You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

EMR集群下Spark超时配置异常:spark.executor.heartbeatInterval需小于等于spark.storage.blockManagerSlaveTimeoutMs问题求助

Hey there! Let's break down why this error is popping up on your EMR cluster and how to fix it.

Why This Error Happens

First, let's clarify the hidden relationship between these Spark parameters:

  • When you don’t explicitly set spark.storage.blockManagerSlaveTimeoutMs, Spark automatically calculates it using your spark.network.timeout value. By default, it takes 50% of the network timeout as the block manager slave timeout.
  • In your case, you set spark.network.timeout=300000 (5 minutes), so the default blockManagerSlaveTimeoutMs becomes 150000 (2.5 minutes). But your spark.executor.heartbeatInterval=200000 (~3.3 minutes) is longer than this derived timeout—this mismatch is exactly what’s triggering the error.

As for the parameter’s origin on EMR: while EMR has some customized Spark defaults, this specific inheritance rule follows core Spark’s behavior. EMR doesn’t override this dependency unless you explicitly set blockManagerSlaveTimeoutMs yourself.

Fixes to Resolve the Issue

You’ve got a few straightforward options to fix this parameter mismatch:

  • Option 1: Adjust the heartbeat interval to fit the default timeout
    Since blockManagerSlaveTimeoutMs is half your spark.network.timeout, set spark.executor.heartbeatInterval to a value ≤ 150000 (e.g., 140000). This aligns with Spark’s recommended ratio (heartbeat interval should be roughly 1/3 to 1/2 of the block manager timeout to avoid false positives).

  • Option 2: Explicitly set blockManagerSlaveTimeoutMs to match your heartbeat interval
    If you need to keep spark.executor.heartbeatInterval=200000, set spark.storage.blockManagerSlaveTimeoutMs to a value ≥ 200000 (e.g., 210000). Just make sure spark.network.timeout is at least twice this value (e.g., 420000)—this ensures the network timeout covers the block manager timeout, preventing other potential issues.

  • Option 3: Align both parameters with Spark’s best practices
    Stick to the recommended ratio: keep spark.executor.heartbeatInterval between 1/3 and 1/2 of spark.network.timeout. For example:

    • spark.network.timeout=300000
    • spark.executor.heartbeatInterval=100000
How to Verify the Fix

After updating your configuration:

  1. Check your Spark application’s startup logs to confirm the actual values of these parameters are being applied.
  2. Or, navigate to the Environment tab in the Spark UI—you’ll see the resolved values for all Spark configuration parameters there.

内容的提问来源于stack exchange,提问作者Neethu Lalitha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:42:42