如何在GCP Spark Scala作业中传递配置文件参数(含Dataproc Web UI)
Hey there! Let's tackle your two questions about working with configuration files for Spark Scala jobs on GCP Dataproc:
1. How to pass configuration file parameters to a Spark Scala job on GCP
There are a couple of reliable approaches depending on whether you want to pass an entire config file or individual parameters:
Option 1: Use a config file stored in GCS
Since Dataproc clusters can natively access Google Cloud Storage (GCS), this is the most common method:
- Upload your config file (e.g.,
app.conf) to a GCS bucket first. - When submitting your job, use the
--filesflag to attach the config file. This copies the file to every worker node's local filesystem:gcloud dataproc jobs submit spark \ --cluster=your-cluster-name \ --jar=gs://your-bucket/path/to/your-job.jar \ --files=gs://your-bucket/path/to/app.conf - In your Scala code, use Spark's
SparkFilesutility to get the local path of the config file, then load it with a library like Typesafe Config:import org.apache.spark.SparkFiles import com.typesafe.config.ConfigFactory import java.io.File val configFilePath = SparkFiles.get("app.conf") val config = ConfigFactory.parseFile(new File(configFilePath)) // Access your parameters val apiEndpoint = config.getString("app.api.endpoint") val batchSize = config.getInt("app.processing.batch-size")
Option 2: Pass individual parameters via --conf
If you don't need a full config file, you can pass key-value pairs directly using the --conf flag (or Dataproc's --properties):
- Submit the job with your custom configs:
gcloud dataproc jobs submit spark \ --cluster=your-cluster-name \ --jar=gs://your-bucket/path/to/your-job.jar \ --conf spark.app.api.endpoint=https://api.example.com \ --conf spark.app.processing.batch-size=1000 - Retrieve them in your Scala code from the SparkConf:
import org.apache.spark.SparkConf import org.apache.spark.sql.SparkSession val sparkConf = new SparkConf() val apiEndpoint = sparkConf.get("spark.app.api.endpoint") val batchSize = sparkConf.getInt("spark.app.processing.batch-size", 500) // Default value if not set val spark = SparkSession.builder().config(sparkConf).getOrCreate()
2. How to pass a config file via the GCP Web UI (Dataproc Console)
If you prefer using the GCP Console instead of the CLI, here's how to attach your config file to a Spark submit command:
- First, upload your config file to a GCS bucket (since the cluster can't access local files from your machine).
- Open the GCP Console's Dataproc page, navigate to your cluster, and click Submit job.
- Fill in the basic job details:
- Job type: Select
Spark - Jar file: Enter the GCS path to your Scala job JAR (e.g.,
gs://your-bucket/jobs/your-app.jar)
- Job type: Select
- Scroll down to the Optional parameters section:
- Under Files, enter the GCS path to your config file (e.g.,
gs://your-bucket/configs/app.conf). This is equivalent to using the--filesflag in the CLI. - If you want to add individual config parameters instead, go to Properties and click Add property. Enter the key (e.g.,
spark.app.api.endpoint) and its corresponding value.
- Under Files, enter the GCS path to your config file (e.g.,
- Fill in any other required fields (like job name) and click Submit to run the job.
内容的提问来源于stack exchange,提问作者Chetan SP
相关产品推荐
相关产品推荐

