Spark JDBC使用200个Executor时,数据库连接数是多少?
Great question—this is a common point of confusion when scaling Spark JDBC jobs, so let’s unpack it step by step:
It’s not about Executor count—it’s about task parallelism
When you configure 200 Executors, that doesn’t directly translate to 200 JDBC connections. Instead, the number of concurrent connections depends on how you’ve set up parallel reading for your JDBC source.Parallel reading = one connection per task
If you’ve enabled parallel JDBC reads (by setting parameters likepartitionColumn,lowerBound,upperBound, andnumPartitions), Spark splits your database query intonumPartitionsseparate tasks. Each task will open its own dedicated JDBC connection to the database, execute its portion of the query, and close the connection once done.
For example: if you setnumPartitions=50, you’ll have up to 50 concurrent JDBC connections at peak, even if you have 200 Executors available (since each Executor can run multiple tasks at once, based on its core count).No parallel configuration = single connection
If you don’t set those partitioning parameters, Spark will run the entire read as a single task. That means only one JDBC connection will be used, regardless of how many Executors you’ve provisioned—most of your Executors will sit idle during this read.A critical note on database limits
Always check your database’s maximum allowed concurrent connections before settingnumPartitions. If you set it higher than the database can handle, you’ll hit connection errors. For example, if your database caps connections at 150, stick tonumPartitions <= 150(or adjust the database’s connection limit if possible).
内容的提问来源于stack exchange,提问作者Atul Verma

