You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在GCP Spark Scala作业中传递配置文件参数(含Dataproc Web UI)

Hey there! Let's tackle your two questions about working with configuration files for Spark Scala jobs on GCP Dataproc:

1. How to pass configuration file parameters to a Spark Scala job on GCP

There are a couple of reliable approaches depending on whether you want to pass an entire config file or individual parameters:

Option 1: Use a config file stored in GCS

Since Dataproc clusters can natively access Google Cloud Storage (GCS), this is the most common method:

  • Upload your config file (e.g., app.conf) to a GCS bucket first.
  • When submitting your job, use the --files flag to attach the config file. This copies the file to every worker node's local filesystem:
    gcloud dataproc jobs submit spark \
      --cluster=your-cluster-name \
      --jar=gs://your-bucket/path/to/your-job.jar \
      --files=gs://your-bucket/path/to/app.conf
    
  • In your Scala code, use Spark's SparkFiles utility to get the local path of the config file, then load it with a library like Typesafe Config:
    import org.apache.spark.SparkFiles
    import com.typesafe.config.ConfigFactory
    import java.io.File
    
    val configFilePath = SparkFiles.get("app.conf")
    val config = ConfigFactory.parseFile(new File(configFilePath))
    
    // Access your parameters
    val apiEndpoint = config.getString("app.api.endpoint")
    val batchSize = config.getInt("app.processing.batch-size")
    

Option 2: Pass individual parameters via --conf

If you don't need a full config file, you can pass key-value pairs directly using the --conf flag (or Dataproc's --properties):

  • Submit the job with your custom configs:
    gcloud dataproc jobs submit spark \
      --cluster=your-cluster-name \
      --jar=gs://your-bucket/path/to/your-job.jar \
      --conf spark.app.api.endpoint=https://api.example.com \
      --conf spark.app.processing.batch-size=1000
    
  • Retrieve them in your Scala code from the SparkConf:
    import org.apache.spark.SparkConf
    import org.apache.spark.sql.SparkSession
    
    val sparkConf = new SparkConf()
    val apiEndpoint = sparkConf.get("spark.app.api.endpoint")
    val batchSize = sparkConf.getInt("spark.app.processing.batch-size", 500) // Default value if not set
    
    val spark = SparkSession.builder().config(sparkConf).getOrCreate()
    
2. How to pass a config file via the GCP Web UI (Dataproc Console)

If you prefer using the GCP Console instead of the CLI, here's how to attach your config file to a Spark submit command:

  • First, upload your config file to a GCS bucket (since the cluster can't access local files from your machine).
  • Open the GCP Console's Dataproc page, navigate to your cluster, and click Submit job.
  • Fill in the basic job details:
    • Job type: Select Spark
    • Jar file: Enter the GCS path to your Scala job JAR (e.g., gs://your-bucket/jobs/your-app.jar)
  • Scroll down to the Optional parameters section:
    • Under Files, enter the GCS path to your config file (e.g., gs://your-bucket/configs/app.conf). This is equivalent to using the --files flag in the CLI.
    • If you want to add individual config parameters instead, go to Properties and click Add property. Enter the key (e.g., spark.app.api.endpoint) and its corresponding value.
  • Fill in any other required fields (like job name) and click Submit to run the job.

内容的提问来源于stack exchange,提问作者Chetan SP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:02:03