Spark 2读取Hive分区表任务数异常问题咨询
Hey there! Let's break down why you're seeing exactly 2000 tasks when reading this table, instead of the smaller number you expected based on your 256MB HDFS block size. Here are the key reasons:
Table partitions ≠ HDFS blocks
Your table is split into 2000 logical Hive partitions, and Spark’s default behavior for reading partitioned Hive tables is to create at least one task per partition. Even if each partition’s underlying Parquet files are way smaller than 256MB (e.g., a few MB each), Spark will still spin up a task for every single partition—hence the 2000 task count matching your partition count exactly.Parquet row groups drive task granularity
Parquet files are organized into row groups, which are the core unit of parallelism for Spark reads. By default, row groups are 128MB (configurable viaparquet.block.size). If each of your 2000 partitions contains exactly one row group (or one small file that fits into a single row group), each will map directly to one Spark task. This aligns perfectly with your observed task count.Input split logic for partitioned tables
When scanning a Hive partitioned table, Spark first enumerates all partitions. For each partition, it generates input splits based on the files in that partition. If a partition’s file is smaller than the HDFS block size (256MB), the entire file becomes one split—and each split translates to one task. With 2000 such partitions, you end up with 2000 tasks.
How to Verify This
You can check the size of files in each partition using this HDFS command:
hdfs dfs -du -h /path/to/your/hive/table/partition=*
If most partitions have files significantly smaller than 256MB, that confirms this is the root cause.
How to Adjust Task Counts
To get task counts closer to your expected total data size / 256MB, you’ll need to reduce the number of small files:
- Use Hive’s
ALTER TABLE your_table CONCATENATE;to merge small Parquet files within partitions. - When writing the table, configure Spark to produce larger files:
- Set
spark.sql.files.maxRecordsPerFileto control how many records go into each file. - Adjust
spark.sql.files.openCostInBytesto make Spark prioritize larger files when creating splits.
- Set
内容的提问来源于stack exchange,提问作者sunil kancharlapalli

