Spark Dataset唯一ID性能对比:row_number与monotonically_increasing_id
row_number() vs monotonically_increasing_id() for Unique Spark IDs Great question—let’s cut through the surface-level explanations and dig into exactly why these two approaches perform so differently under the hood.
1. row_number() with an Unpartitioned Window: The Heavyweight Contender
When you use row_number().over(Window.orderBy("a column")) without specifying a partitionBy clause, you’re asking Spark to do two very expensive things:
- Global Shuffle & Sort: Spark needs to rearrange your entire dataset across the cluster so that all rows are ordered by your chosen column. This triggers a full shuffle operation—data is serialized, sent over the network between nodes, and written to disk temporarily. Shuffles are one of the biggest performance bottlenecks in Spark because they’re I/O and network-heavy.
- O(n log n) Time Complexity: Sorting large datasets is computationally expensive. For a dataset of size
n, global sorting takes O(n log n) time, which scales poorly as your data grows. Even worse, if yourorderBycolumn has skewed values (e.g., a lot of rows share the same value), some shuffle partitions will end up with way more data than others, leading to straggler tasks that slow down the entire job. - Memory & Disk Pressure: Spark has to hold sorted chunks of data in memory to generate continuous row numbers. If your dataset is too large to fit in memory, Spark will spill to disk, adding even more I/O overhead.
In short, this approach forces Spark to do a ton of work just to assign IDs—work that’s completely unnecessary if you don’t care about sequential IDs.
2. monotonically_increasing_id(): The Lightweight, Distributed Workhorse
This function is built for speed because it leverages Spark’s distributed nature without any cross-node coordination:
- No Shuffle, No Sort: Each task in Spark processes a single partition of data.
monotonically_increasing_id()generates IDs locally within each task using a simple counter. The ID is a 64-bit integer where the high 31 bits represent the task’s unique ID, and the low 33 bits represent the row index within the task. This means every task can generate IDs independently—no data needs to be sent between nodes, no sorting is required. - O(n) Time Complexity: Each partition is processed in linear time, just counting rows as they pass through. There’s no expensive sorting step, so this scales beautifully even for massive datasets.
- Minimal Overhead: The function only needs a tiny amount of memory to track the row counter per task. There’s no disk spillage, no network traffic, no waiting on straggler tasks.
The tradeoff is non-sequential IDs, but since you’ve stated that doesn’t matter, this is a no-brainer for performance.
Real-World Performance Comparison
- Small Datasets: You might not notice a huge difference—both methods will run quickly. But even here,
monotonically_increasing_id()will edge out the window approach by avoiding unnecessary shuffle/sort steps. - Large Datasets: The gap becomes massive. For TB-scale data,
row_number()could take hours (or fail entirely due to memory issues), whilemonotonically_increasing_id()will finish in minutes, limited only by how fast your cluster can read the input data. - Skewed Data: If your
orderBycolumn has significant skew,row_number()will crawl to a halt as one or more tasks get stuck processing giant partitions.monotonically_increasing_id()doesn’t care about data distribution—it’ll zip through regardless.
Final Recommendation
If sequential IDs aren’t a requirement, always use monotonically_increasing_id(). It’s orders of magnitude faster, more scalable, and far less likely to run into performance or stability issues. The row_number() approach should only be used when you absolutely need continuous, ordered IDs—and even then, try to add a partitionBy clause to break the data into smaller chunks and reduce the sorting overhead.
内容的提问来源于stack exchange,提问作者Henrique Goulart

