Windows命令行运行PySpark失败求助(Spark-shell可正常运行)
Hey there! Since your spark-shell runs fine, your core Spark and Java setup is mostly solid—let’s focus on the PySpark-specific quirks that might be tripping you up. Here’s a step-by-step breakdown to diagnose the issue:
Verify your Python environment alignment
Since you installed PySpark via Anaconda’s pip, make sure thepysparkcommand is using your Anaconda Python instance. Run these commands in your command prompt to check:where python: Confirm the first path listed points to your Anaconda Python (e.g.,C:\Users\heyde\Anaconda3\python.exe).python --version: Ensure your Python version is compatible with your Spark version (Spark 2.x supports Python 2.7/3.4+, Spark 3.x requires 3.6+).- Try importing PySpark directly in a Python shell:
If this throws an error, it points to an issue with the PySpark package itself or Python’s ability to locate it.import pyspark from pyspark.sql import SparkSession spark = SparkSession.builder.getOrCreate()
Double-check SPARK_HOME environment variable
You setSPARK_HOMEtoC:\Users\heyde\Anaconda3\Lib\site-packages\pyspark—that’s correct for a pip-installed PySpark. But confirm it’s actually active:- Run
echo %SPARK_HOME%in your command prompt. If the output doesn’t match your intended path, verify you added it to the system environment variables (not just user variables) and restart your command prompt to apply changes.
- Run
Try launching PySpark via the full path or Anaconda Prompt
Sometimes the defaultpysparkcommand might not pick up your Anaconda environment properly. Test these alternatives:- Run the full path to the PySpark script:
C:\Users\heyde\Anaconda3\Scripts\pyspark.cmd - Open the Anaconda Prompt (instead of regular Command Prompt) and run
pyspark—this auto-loads Anaconda’s environment variables, which often resolves path conflicts.
- Run the full path to the PySpark script:
Check for missing winutils.exe (a common Windows gotcha)
Spark relies on Hadoop utilities to work on Windows, and pip-installed PySpark doesn’t includewinutils.exeby default. If you’re seeing errors related to Hadoop or file system operations:- Download
winutils.exematching your Spark’s Hadoop version (e.g., Spark 2.4.x uses Hadoop 2.7, Spark 3.x uses Hadoop 3.2). - Place it in
%SPARK_HOME%\bin. - Set a
HADOOP_HOMEenvironment variable to yourSPARK_HOMEpath (e.g.,C:\Users\heyde\Anaconda3\Lib\site-packages\pyspark), and add%HADOOP_HOME%\binto your system PATH.
- Download
Validate JAVA_HOME details
You’ve setJAVA_HOMEtoC:Java, but let’s confirm it’s pointing to the right JDK:- Run
java -versionandjavac -version—both should output1.8.0_161. If not, there might be another JDK installed that’s overriding yourJAVA_HOMEsetting. EnsureC:Java\binis at the top of your system PATH to prioritize it.
- Run
If you run into specific error messages during these steps, sharing those details will help narrow things down even further!
内容的提问来源于stack exchange,提问作者Michał Heydel

