You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Dataproc集群缩容请求超时未完成,能否重新提交该请求?

Dataproc Cluster Resize Stuck at "cluster update in progress"

Let me break this down based on my hands-on experience with Dataproc:

Is this normal?

In most cases, scaling down a Dataproc cluster should wrap up within minutes to 15-20 minutes—waiting an hour is way longer than the expected timeframe, so this isn’t "normal" behavior. That said, there are a few common scenarios that can cause delays:

  • Active jobs running: If there are ongoing tasks on the cluster, Dataproc will often wait for jobs to finish or migrate tasks to remaining workers before terminating nodes. This can drag on depending on the job’s complexity and data volume.
  • Resource cleanup bottlenecks: Terminating worker nodes involves uninstalling Dataproc components, syncing metadata, and releasing underlying GCE resources (like persistent disks or network interfaces). Rarely, this cleanup process gets stuck due to transient infrastructure issues.
  • Cluster state inconsistencies: Glitches with the control plane or node agents can slow down the update process, leaving the operation in a pending state.

Can I resubmit the resize request?

Don’t do this. Submitting another resize request while an existing update is in progress will create conflicting operations. This can lead to messy outcomes—like partial scaling, permanently stuck nodes, or even making the cluster unmanageable. You need to let the current operation resolve (or troubleshoot why it’s stuck) before making any further changes.

Troubleshooting steps to diagnose the issue

  1. Check the stuck operation’s details:
    Use the gcloud CLI to list ongoing operations for your cluster:

    gcloud dataproc operations list --cluster=YOUR_CLUSTER_NAME
    

    Grab the operation ID of the stuck resize, then pull detailed logs to see where it’s hung:

    gcloud dataproc operations describe OPERATION_ID
    

    This will show you the current progress, any error messages, or which step the operation is stuck on.

  2. Verify active jobs:
    Check if there are any running or pending jobs on the cluster that might be blocking the resize:

    gcloud dataproc jobs list --cluster=YOUR_CLUSTER_NAME
    

    By default, Dataproc won’t terminate workers mid-job, so you’ll need to wait for jobs to finish or manually terminate them (if you’re okay with potential data loss).

  3. Inspect node status:
    Check the state of your cluster’s nodes in GCE to see if any workers are stuck terminating:

    gcloud compute instances list --filter="name:YOUR_CLUSTER_NAME"
    

    Nodes stuck in a TERMINATING state for an hour usually point to an issue with underlying GCE resource deletion.

  4. Contact support if needed:
    If the operation is stuck for more than 2 hours, or the logs show unclear errors, reach out to GCP support with your cluster name, operation ID, and a description of the issue for deeper troubleshooting.


内容的提问来源于stack exchange,提问作者K2J

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:33:24