Spark中关联大表与小表时,小表大字段是否会实际重复存储?
Great question—this is a common concern when dealing with large payloads in joins, especially with asymmetric table sizes like yours. Let’s break down how Spark handles this:
1. 优先触发的Broadcast Join:几乎不会重复存储z2
Since T2 is much smaller than T1, Spark will almost certainly use a Broadcast Join (you can confirm this by checking the query execution plan with EXPLAIN). Here’s what happens:
- T2 is serialized once and broadcast to every executor node in your cluster. Each executor stores a single copy of the entire T2 dataset (including all z2 values) in memory (or disk if memory is tight).
- When processing partitions of T1, Spark just references the local broadcast copy of T2 to match rows on
x2 = z1. The z2 values aren’t duplicated in memory for each matching T1 row—they’re just looked up from the single broadcast copy when building the result rows.
This is the most efficient scenario here, and it completely avoids redundant in-memory storage of the large z2 values.
2. Shuffle Join(极少出现的情况):依然避免冗余存储
If for some reason Spark doesn’t use Broadcast Join (e.g., T2 is just over the broadcast threshold), it will fall back to a Shuffle Join. Even here:
- T2 is shuffled by
z1, so each executor gets a subset of T2 rows grouped by theirz1value. For each uniquez1, the corresponding z2 is stored only once per executor partition. - T1 is shuffled by
x2to match T2’s partitions, and when joining, Spark again references the single stored z2 value for each matching group of T1 rows. No duplicate z2 copies are created during the join processing.
3. 结果持久化:会物理重复存储(不可避免的业务需求)
The only time you’ll see duplicate z2 values is when you write the query result to a persistent storage system (like HDFS, S3, or a database). Since the result set requires each row to contain both x1 and z2, every row in the output will have the full z2 value for its matching group.
If this storage overhead is a problem, you could consider:
- Storing only
x1and the join keyx2in the result, then joining back to T2 only when you need to access z2 later. - Compressing the output data (Spark supports various compression codecs like Snappy or Gzip) to reduce the physical storage footprint of the large z2 values.
内容的提问来源于stack exchange,提问作者Vitaliy

