Apache Spark节点为何需要访问数据文件路径?
Great question—this trips up a lot of new Spark users, so let's break down exactly what's happening here.
First, let's clear up a common misunderstanding: when you run distFile = sc.textFile("data.txt"), you're not loading the entire file into the Driver and then sending the full dataset to Worker nodes. All this line does is create a logical RDD definition—it tells Spark "I want to create an RDD from this file", but no actual data is read or transferred yet. Spark uses lazy execution, so data loading only happens when you trigger an action (like distFile.count() or distFile.collect()).
Here's what happens once you run an action:
- The Driver splits the input file into multiple partitions (based on file size and Spark's default partition settings).
- The Driver assigns each partition's computation task to a Worker node.
- Each Worker node is responsible for reading its assigned partition directly from the source file.
This is where the "file must exist on all Workers" rule comes in: if the file is only on the Driver's local filesystem, the Worker nodes have no way to access it when they try to read their partition. They can't reach the Driver's local disk, so the task would fail with a file-not-found error.
This design is intentional—it's part of Spark's focus on data locality. Instead of moving large datasets across the network to where the computation happens, Spark sends the computation tasks to where the data lives. This drastically reduces network overhead and improves performance. For example:
- If your
data.txtis 1GB and split into 2 partitions, having Workers read local copies means each Worker only reads 500MB locally (fast, no network lag). - If the Driver had to send the full 1GB to Workers, you'd waste time and bandwidth transferring the entire dataset first.
That's why the two solutions work:
- Copying the file to the same path on all Workers: Each Worker can read its partition from its own local disk, maximizing speed.
- Using a network-mounted shared filesystem (like NFS or HDFS): All nodes access the same shared storage via the same path, so Workers can pull their partitions directly from the shared source without local copies.
To sum it up: The Driver doesn't distribute the RDD data—it distributes the instructions to read the data. Workers need access to the file path because they're the ones doing the actual data reading when tasks run.
内容的提问来源于stack exchange,提问作者toto'

