如何在Mesos下无需重部署实现Apache Spark运行时资源扩容?
Great question! I’ve dealt with exactly this scenario when running Spark on DC/OS with the SMACK stack, so let’s walk through the practical ways to adjust spark.cores.max at runtime without restarting your application.
First, a quick recap: You’re right that spark.cores.max caps the total CPU cores your Spark app can request across the cluster (in Mesos coarse-grained mode, which I assume you’re using since you mentioned Marathon for scaling Kafka/Cassandra). Spark will auto-launch new executors as needed only if it hasn’t hit this limit—so the key is to update this value while the app is running.
Here are the most reliable methods:
1. Use Spark’s Application REST API
Spark’s built-in REST API lets you modify runtime configuration for running apps, including spark.cores.max. This works for both Standalone and Mesos coarse-grained clusters:
- Find your app’s REST endpoint: On DC/OS, you can get the Spark app’s API URL from the DC/OS UI (under the Spark service’s details) or via the DC/OS CLI:
dcos spark info --app-id <your-spark-app-id> - Update the core limit: Send a POST request to the cores endpoint with your new max value:
For example, if you want to bump from 8 cores to 16:curl -X POST http://<spark-api-url>/applications/<app-id>/cores -d "cores=<new-max-value>"curl -X POST http://spark.marathon.l4lb.thisdcos.directory:6066/applications/my-spark-app/cores -d "cores=16"
2. Use the DC/OS Spark CLI
DC/OS has a dedicated Spark CLI that simplifies updating running apps. This is my go-to method because it integrates directly with DC/OS’s resource management:
dcos spark update --app-id <your-spark-app-id> --conf spark.cores.max=<new-max-value>
This command tells the DC/OS Spark framework to adjust the core limit for your running application, and Spark will immediately start using the new cap when scheduling future tasks/executors.
3. Programmatic Adjustment (For Custom Apps)
If your app has a custom control plane (like an Akka-based service in your SMACK stack), you can use the SparkContext’s conf object to update the setting programmatically—though note this only works for certain dynamic configs:
// In your Spark application code spark.sparkContext.conf.set("spark.cores.max", "16")
However, this requires your app to expose a way to trigger this change (e.g., an HTTP endpoint or message queue listener). It’s more work but useful if you want to automate scaling based on cluster metrics.
Important Notes & Limitations
- Coarse-grained mode only: These methods only work in Mesos coarse-grained sharing mode (which you’re already using, since you mentioned
spark.cores.max). Fine-grained mode doesn’t support runtime core scaling. - Resource availability: Spark can only use the new core limit if DC/OS has idle CPU cores available in the cluster. Check the DC/OS UI’s cluster overview to confirm free resources.
- Executor scaling timing: Spark won’t immediately spin up new executors—they’ll be launched as needed when new tasks are scheduled (e.g., for the next batch in a Streaming app, or new stages in a batch job).
- Verify changes: Use the Spark UI (accessible via DC/OS) or the REST API to confirm the new
spark.cores.maxvalue:curl http://<spark-api-url>/applications/<app-id>/conf | grep spark.cores.max
内容的提问来源于stack exchange,提问作者Benedikt Wedenik

