PySpark深度学习报错:AttributeError: module 'tensorflow' has no attribute 'Session'
PySpark深度学习运行失败问题解决
问题重现
在PySpark中构建深度学习图像分类流水线时执行失败,代码如下:
featurizer = DeepImageFeaturizer(inputCol="image", outputCol="features", modelName="InceptionV3") lr = LogisticRegression(maxIter=5, regParam=0.03, elasticNetParam=0.5, labelCol="label") sparkdn = Pipeline(stages=[featurizer, lr]) spark_model = sparkdn.fit(train)
报错信息
File ~/anaconda3/lib/python3.9/site-packages/sparkdl/transformers/keras_applications.py:45, in KerasApplicationModel.getModelData(self, featurize) 44 def getModelData(self, featurize): ---> 45 sess = tf.Session() 46 with sess.as_default(): 47 K.set_learning_phase(0) AttributeError: module 'tensorflow' has no attribute 'Session'
已尝试操作
- 重新安装TensorFlow
- 启动PySpark时指定依赖包:
pyspark --packages databricks:spark-deep-learning:0.1.0-spark2.1-s_2.11
问题原因
你使用的spark-deep-learning 0.1.0版本是针对Spark 2.1和TensorFlow 1.x开发的,而当前环境安装的是TensorFlow 2.x版本。TensorFlow 2.x中已经移除了tf.Session()接口,改用即时执行模式,因此老版本的spark-deep-learning代码无法在TF2.x环境中运行。
解决方案
方案1:降级TensorFlow到1.x版本
安装与spark-deep-learning 0.1.0兼容的TensorFlow 1.x版本,注意匹配Python版本(Python 3.9需选择支持该版本的TF1.x分支,若版本不匹配建议选择其他方案):
pip install tensorflow==1.15.5
方案2:启用TensorFlow 2.x兼容模式
在代码开头添加以下代码,让TF2.x模拟TF1.x的运行环境:
import tensorflow.compat.v1 as tf tf.disable_v2_behavior()
该方法无需降级TF,但可能存在部分API兼容问题,需自行测试验证。
方案3:升级spark-deep-learning版本
若你的Spark版本高于2.1,建议升级到适配Spark 3.x和TensorFlow 2.x的spark-deep-learning版本,比如使用针对Spark 3.0的版本:
pyspark --packages databricks:spark-deep-learning:1.5.0-spark3.0-s_2.12
升级后可直接使用TensorFlow 2.x版本,避免版本兼容问题。
内容的提问来源于stack exchange,提问作者SHR_Az
相关产品推荐
相关产品推荐

