能否以集群模式运行Dataproc作业?GCP Dataproc部署模式咨询
Absolutely, GCP Dataproc does support the spark.submit.deployMode=cluster mode when submitting PySpark jobs via the gcloud dataproc jobs submit pyspark command!
By default, Dataproc uses client mode for these submissions—meaning the Spark driver runs on the machine where you execute the gcloud command. But switching to cluster mode is straightforward: you just need to pass the spark.submit.deployMode property via the --properties flag.
Here's a concrete example of the command to submit a PySpark job in cluster mode:
gcloud dataproc jobs submit pyspark gs://your-bucket-path/your-spark-script.py \ --cluster=your-dataproc-cluster-name \ --region=your-gcp-region \ --properties spark.submit.deployMode=cluster
A quick note on why you might want this: when using cluster mode, the Spark driver runs on one of the cluster's worker nodes instead of your local machine. This is great for long-running jobs, large workloads, or cases where your local machine has limited resources (since you don't have to keep the terminal open or worry about local network drops killing the driver).
内容的提问来源于stack exchange,提问作者jamiet

