无需HDFS,Spark如何基于各Slave节点的分段本地文件运行任务?
Got it, let's tackle this problem head-on. The core issue here is that your current code only looks for data on the Master node's local filesystem, while your segmented data lives on the Slave nodes. Spark doesn't automatically scan Slave local paths unless you explicitly point it to them. Here are two straightforward ways to fix this without relying on HDFS or external tools:
1. Explicitly List All Slave Node Local Paths
Since each Slave holds a portion of your data, you can tell Spark to read from all their local paths directly using the file:// protocol.
First, you'll need to know the hostnames/IPs of your Slave nodes (or use identical relative paths if all Slaves share the same local path structure). Then pass all these paths to the sequenceFile method—Spark will distribute the read tasks to the corresponding Slaves, where Executors will pull the local data.
Here's how to adjust your code:
import java.util.Arrays; import java.util.List; // Replace these paths with your actual Slave local paths (use file:// protocol) List<String> slaveDataPaths = Arrays.asList( "file:///path/to/segmented/data/on/slave1.seq", "file:///path/to/segmented/data/on/slave2.seq" ); // Pass the comma-separated paths to sequenceFile JavaPairRDD<IntWritable, VectorWritable> seqVectors = sc.sequenceFile(String.join(",", slaveDataPaths), IntWritable.class, VectorWritable.class);
Key Notes:
- Ensure each Slave's local file is accessible to the Executor process running on that node (check file permissions—Executor user needs read access).
- If all Slaves use the exact same local path (e.g.,
/data/part.seq), you can even just pass that singlefile:///data/part.seqpath. Spark will automatically have each Executor read the local file on its own Slave node. This works because thefile://path without a hostname resolves to the local filesystem of the node where the Executor is running.
2. Use a Custom Input Format (Advanced)
If you need more control over how data is split across Slaves, you can create a custom InputFormat that detects and reads segmented local data. However, this is overkill for most cases—method 1 is simpler and should cover your needs.
Why Your Original Code Failed
To recap: When you pass a single local path to sequenceFile, Spark's Driver (running on Master) tries to read that path from the Master's local filesystem. Since Master has no data, you end up with an empty RDD—hence no results, no errors (empty RDDs are valid in Spark). The Slave nodes' data was never accessed because Spark wasn't instructed to look there.
内容的提问来源于stack exchange,提问作者user3086871

