Windows下Python3.14中PySpark调用.show()报Python Worker意外退出错误
Windows 10下PySpark调用DataFrame.show()触发Python Worker崩溃问题
在Windows 10系统中使用PySpark时,调用DataFrame的.show()方法出现错误,任务因Python Worker意外崩溃而失败。
环境信息
- 操作系统:Windows 10
- Spark:Apache Spark(PySpark)
- IDE:VS Code
- Python版本:3.14
错误表现
调用.show()方法时,Spark任务报错,提示Python Worker意外崩溃,导致任务执行失败(原问题附带3张错误截图)。
复现代码
from pyspark.sql import SparkSession from pyspark.sql.types import StructType, StructField, StringType spark = SparkSession.builder \ .appName("Test") \ .master("local[*]") \ .getOrCreate() emp_schema = StructType([ StructField("employee_id", StringType(), True), StructField("department_id", StringType(), True), StructField("name", StringType(), True), StructField("age", StringType(), True), StructField("gender", StringType(), True), StructField("salary", StringType(), True), StructField("hire_date", StringType(), True) ]) emp_data = [ ["001", "101", "John Doe", "30", "Male", "50000", "2015-01-01"], ["002", "101", "Jane Smith", "25", "Female", "45000", "2016-02-15"], ["003", "102", "Bob Brown", "35", "Male", "55000", "2014-05-01"], ["004", "102", "Alice Lee", "28", "Female", "48000", "2017-09-30"], ["005", "103", "Jack Chan", "40", "Male", "60000", "2013-04-01"], ["006", "103", "Jill Wong", "32", "Female", "52000", "2018-07-01"], ["007", "101", "James Johnson", "42", "Male", "70000", "2012-03-15"], ["008", "102", "Kate Kim", "29", "Female", "51000", "2019-10-01"], ["009", "103", "Tom Tan", "33", "Male", "58000", "2016-06-01"], ["010", "104", "Lisa Lee", "27", "Female", "47000", "2018-08-01"], ["011", "104", "David Park", "38", "Male", "65000", "2015-11-01"], ["012", "105", "Susan Chen", "31", "Female", "54000", "2017-02-15"], ["013", "106", "Brian Kim", "45", "Male", "75000", "2011-07-01"], ["014", "107", "Emily Lee", "26", "Female", "46000", "2019-01-01"], ["015", "106", "Michael Lee", "37", "Male", "63000", "2014-09-30"], ["016", "107", "Kelly Zhang", "30", "Female", "49000", "2018-04-01"], ["017", "105", "George Wang", "34", "Male", "57000", "2016-03-15"], ["018", "104", "Nancy Liu", "29", "Female", "50000", "2017-06-01"], ["019", "103", "Steven Chen", "36", "Male", "62000", "2015-08-01"], ["020", "102", "Grace Kim", "32", "Female", "53000", "2018-11-01"] ] # 冗余导入可删除 # from pyspark.sql.types import StructType, StructField, StringType emp_schema = "employee_id string, department_id string, name string, age string, gender string, salary string, hire_date string" emp = spark.createDataFrame(emp_data, emp_schema) emp.show()
解决方案
1. 匹配Python与Spark版本兼容性
Python 3.14属于极新版本,目前Spark 3.5.x及以下版本均未提供官方支持。建议将Python降级到3.8~3.11区间,这是Spark官方明确兼容的版本范围。
2. 配置Windows环境变量
- 添加
PYSPARK_PYTHON环境变量,值为你的Python可执行文件绝对路径(如C:\Python310\python.exe) - 添加
PYSPARK_DRIVER_PYTHON环境变量,值与上述路径一致,确保Driver和Worker进程使用同一Python解释器 - 确保Spark安装目录下的
bin文件夹已添加到系统PATH中
3. 显式指定SparkSession的Python路径
在创建SparkSession时,直接配置Python路径参数:
spark = SparkSession.builder \ .appName("Test") \ .master("local[*]") \ .config("spark.pyspark.python", "C:/Python310/python.exe") \ .config("spark.pyspark.driver.python", "C:/Python310/python.exe") \ .getOrCreate()
4. 排查安全软件拦截
Windows防火墙或杀毒软件可能会拦截Spark的Python Worker进程,可临时关闭相关软件测试是否恢复正常。
内容的提问来源于stack exchange,提问作者Deepika Goyal
相关产品推荐
相关产品推荐

