Spark Pipe函数抛出No such file or directory错误求助
Hey there, let's work through that frustrating "No such file or directory" error you're seeing when using Spark Pipe on your EMR cluster. Here are the most common fixes to get your script running smoothly:
1. Use SparkFiles to Reference the Script Path on Executors
When you use sc.addFile(), Spark copies your script to a temporary directory on each executor node—not the original /home/hadoop/ path you specified. If you try to call the script using the original path in your pipe() call, executors won't find it.
Instead, import org.apache.spark.SparkFiles and use SparkFiles.get() to fetch the correct path on each executor. Here's your modified code:
import org.apache.spark._ import org.apache.spark.SparkFiles // Don't forget this import! val distScript = "/home/hadoop/PipeEx.sh" val distScriptName = "PipeEx.sh" sc.addFile(distScript) val ipData = sc.parallelize(List("asd","xyz","zxcz","sdfs")) // Use SparkFiles to get the correct script path on executors val result = ipData.pipe(SparkFiles.get(distScriptName)).collect() result.foreach(println)
2. Ensure Your Script Has Executable Permissions
Even if the script is copied to executors, it won't run if it doesn't have execute permissions. On your EMR master node, run this command to fix the permissions:
chmod +x /home/hadoop/PipeEx.sh
3. Add a Shebang Line to Your Script
Make sure your PipeEx.sh script starts with a shebang line that specifies the shell interpreter. Without this, the system won't know how to execute the script. For example:
#!/bin/bash # Rest of your script code here echo "Processing input: $1"
4. Verify the Script Exists on the Master Node First
Double-check that PipeEx.sh is actually present at /home/hadoop/ on your EMR master node. Run this command to confirm:
ls -l /home/hadoop/PipeEx.sh
Try these steps in order—most likely the first two fixes will resolve your error.
内容的提问来源于stack exchange,提问作者Kask

