You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在运行中的Spark作业中优雅移除Worker节点?

Gracefully Removing a Spark Worker Node from a Running Cluster

Great question! Yes, you can safely take a Spark Worker node offline without disrupting in-flight jobs. The key is to tell the scheduler to stop assigning new tasks to the Worker, wait for its current workload to finish, then shut it down cleanly. Here's how to pull this off:

Step 1: Mark the Worker for Decommissioning

This step signals to the Spark Master that the Worker should no longer receive new tasks, while letting it finish all currently running jobs. You can do this in two ways:

Via the Spark Master UI

  1. Open the Spark Master Web UI (default port 8080) in your browser.
  2. Locate the target Worker node in the "Workers" list.
  3. Click the Worker's ID to open its detail page.
  4. Look for the Decommission button and click it. The Master will immediately stop sending new tasks to this Worker.

Via Command Line (REST API)

If you prefer scripted/headless operations, use curl to call the Master's REST endpoint:

curl -X POST http://<master-host>:<master-port>/v1/workers/<worker-id>/decommission

You can find the Worker ID in the Master UI's Workers list.

Step 2: Wait for the Worker to Complete Its Current Tasks

After marking the Worker for decommission, leave it running until all its active tasks finish. You can monitor progress via:

  • The Master UI: Check the "Running Tasks" count for the Worker—wait until it drops to 0.
  • The Worker's own UI (default port 8081): Verify that no tasks are still in progress.

If you've enabled storage decommissioning (via spark.storage.decommission.enabled=true), the Worker will also migrate any cached RDD data to other available nodes during this phase. Make sure your cluster has enough spare capacity to handle this migration to avoid data loss.

Step 3: Stop the Worker Process

Once the Worker has no running tasks left, shut it down cleanly:

On the Worker Node

If you're using Spark's built-in scripts:

cd <spark-install-directory>/sbin
./stop-worker.sh

If Using Systemd (for managed clusters)

If the Spark Worker is running as a systemd service:

sudo systemctl stop spark-worker

Notes for Older Spark Versions (Pre-3.0)

The formal decommissioning feature was added in Spark 3.0. For earlier versions (2.x):

  1. First, monitor the Worker's task list until all running jobs complete.
  2. Once tasks are done, go to the Master UI and click Remove Worker to take it out of the cluster's available node pool.
  3. Finally, stop the Worker process using the same commands above.
    Be careful: Pre-3.0, there's no built-in way to block new task assignments, so you'll need to wait for existing tasks to finish before removing the Worker to avoid interruptions.

By following these steps, you can safely take a Worker offline without breaking your Spark jobs. Just be patient and let the node wrap up its current work before shutting it down!

内容的提问来源于stack exchange,提问作者sam93

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:02:40