单Spark Session通过Akka HTTP并行执行SQL时多Job创建且多数被跳过
Hey there! Let's dig into why you're seeing multiple skipped Job IDs in the Spark UI when running parallel queries through your Akka HTTP API. Here are the most likely culprits and how to fix them:
1. Akka Dispatcher Thread Pool vs. Spark's Resource Limits
First up, your Akka Dispatcher configuration might be spawning more concurrent requests than your Spark cluster can handle. Spark can only process a limited number of jobs in parallel based on your cluster's resources (executor cores, memory, etc.). If your Dispatcher's thread count exceeds this limit:
- Spark will queue excess jobs
- If requests time out or are canceled (e.g., by Akka HTTP's default request timeout settings), those queued jobs get marked as "skipped"
Fix:
Tune your Dispatcher to match Spark's capacity. For example, use a thread pool sized to your Spark cluster's parallelism:
akka.actor.spark-query-dispatcher { type = Dispatcher executor = "thread-pool-executor" thread-pool-executor { core-pool-size-min = 2 core-pool-size-max = 4 # Match to spark.default.parallelism or executor core count } throughput = 50 }
Make sure your query handling actors/routes use this dedicated dispatcher instead of the default one.
2. Accidental Multiple Spark Sessions
If your Spark Service isn't maintaining a single global Spark Session instance, each request might be creating a new session. Multiple sessions will compete fiercely for cluster resources, leading to job queuing and skipped jobs as resources get exhausted.
Fix:
Verify your Spark Session is initialized once (e.g., as a singleton object in Scala) and reused across all requests. Avoid creating a new session per query—this is a common pitfall that wreaks havoc with parallel execution.
3. Spark's Query Reuse & "Skipped" Jobs (It Might Be an Optimization!)
Sometimes Spark marks jobs as "skipped" when it can reuse results from a previously executed query. For example:
- If two identical queries are submitted in parallel
- If a query's dependent stages have already completed successfully from another job
This is actually a Spark optimization, not an error! Check the Spark UI's Job details: look for "Skipped" stages and see if they note "already computed" or similar messages.
Fix:
If this is the case, you don't need to fix anything—Spark is just being efficient. But if you didn't intend to run duplicate queries, check your API client or route logic for accidental repeated requests.
4. Request Timeouts & Canceled Jobs
Akka HTTP has default request timeout settings. If a query takes longer than the timeout to start executing (because it's stuck in Spark's job queue), Akka might cancel the request, which in turn cancels the Spark job—marking it as skipped.
Fix:
Adjust Akka's request timeout for query routes to match your expected query duration:
akka.http.server.request-timeout = 5m # Or longer for complex, data-heavy queries
Alternatively, implement asynchronous query tracking (e.g., return a job ID immediately, then let clients poll for results) instead of blocking on the query completion.
5. Spark Configuration Misalignment
Double-check these Spark settings to ensure they support parallel job execution:
spark.default.parallelism: Controls the default number of partitions for RDDs (should match your cluster's core count)spark.sql.shuffle.partitions: Controls shuffle partitions for SQL queries (adjust based on data size and available cores)spark.scheduler.mode: Ensure it's set toFAIR(default isFIFO) if you want to prioritize multiple jobs fairly instead of running them in strict sequence.
Next Steps to Diagnose
- Open the Spark UI and click on a "skipped" job to view its details. Look for cancellation reasons or skipped stage explanations—this will give you the clearest clue.
- Enable debug logging for your Akka HTTP routes and Spark service to see if requests are being duplicated or canceled prematurely.
- Test with a single query first to confirm it runs without skipped jobs, then gradually add parallel requests to isolate the issue.
内容的提问来源于stack exchange,提问作者Rajat Mishra

