无法通过Livy向AWS EMR上的Spark传递Python依赖求助
Let's break down exactly what's happening here and fix it:
The Root Cause
You're hitting this error because of a redundant (and incorrectly targeted) sc.addFile() call. When you include dependencies.zip in Livy's pyFiles parameter, Livy automatically handles:
- Downloading the zip from S3 to all Spark nodes' temporary directories
- Adding the zip to the Python
PYTHONPATHso your code can import dependencies like pandas or sqlalchemy
But when you call sc.addFile('dependencies.zip') in global_settings.py, Spark looks for this file in the local filesystem of the driver node (not the distributed path managed by Livy/Spark), hence the FileNotFoundException.
Fixes (Ordered by Recommendation)
1. Remove the sc.addFile() call (Best Option)
Simply delete this line from global_settings.py:
sc.addFile('dependencies.zip')
Livy's pyFiles parameter already does all the heavy lifting for distributing and configuring your dependency zip. Your code will be able to import pandas, sqlalchemy, etc., without any extra work here.
2. Use the full S3 path if you must keep sc.addFile()
If you have a specific reason to manually add the file (though this is unnecessary here), you need to pass the complete S3 URL instead of just the filename:
sc.addFile("s3://my-bucket/src/dependencies.zip")
This tells Spark to pull the file directly from S3 instead of looking locally. Note that this will duplicate the distribution work already done by Livy, so it's not ideal.
Verify the Fix
To confirm your dependencies are loading correctly, add a quick check at the start of my_job.py:
import sys import pandas print("Current PYTHONPATH:", sys.path) print("Successfully imported pandas version:", pandas.__version__)
This will show you if the zip is in your path and if the dependency is usable.
Why the other .py files work but the zip doesn't
Livy places .py files from pyFiles directly into the job's working directory (the /mnt/tmp/... path you see in the logs), so importing them works as expected. Zip files, however, are stored in a separate managed directory by Spark and added to PYTHONPATH without being moved to the working directory—hence why referencing the zip by filename alone fails.
内容的提问来源于stack exchange,提问作者Z.Wei

