spark-sklearn 0.2.3中GridSearchCV的best_score_失效,如何安装旧版本?
First off, I notice you might have used an incorrect pip command when trying to install the old version—instead of pip install spark-sklearn-0.2.0, you need to specify the version using == with the official package name. That's probably why your initial install failed. Let's walk through the correct steps for both fixing the installation and deploying it across your cluster.
Step 1: Fix the Basic Installation Command
The right command to install spark-sklearn 0.2.0 is:
pip install spark-sklearn==0.2.0
If you run into permission issues (common on shared clusters), add the --user flag to install it in your user directory:
pip install --user spark-sklearn==0.2.0
If there are dependency conflicts with existing packages, you can force a reinstall (use cautiously, as it may override other versions):
pip install --force-reinstall spark-sklearn==0.2.0
Step 2: Install Across Your Cluster
How you deploy this to all nodes depends on your cluster setup:
Option 1: Manual Installation (Small Clusters)
If you have a small number of nodes, log into each one via SSH and run the pip command above. Make sure you're using the same Python environment that your Spark jobs will use (e.g., if you're using a virtualenv or conda environment, activate it first before installing).
Option 2: Batch Installation via Cluster Management Tools (Large Clusters)
For larger clusters, manual installation isn't feasible. Here are some common approaches:
- Conda Environments: If your cluster uses Conda for package management, create or update a shared environment with the specified version, then sync it across all nodes:
Consult your cluster docs on how to propagate Conda environments to all worker nodes.conda install -n your-spark-env spark-sklearn=0.2.0 - Spark
--py-filesFlag: Download the spark-sklearn 0.2.0 wheel file first, then include it when submitting your Spark job. This will distribute the package to all worker nodes dynamically:spark-submit --py-files spark_sklearn-0.2.0-py2.py3-none-any.whl your_job_script.py - Bootstrap Scripts: If your cluster supports bootstrap scripts (like AWS EMR, Databricks, or on-prem clusters with orchestration tools), write a simple bash script containing the pip install command, and configure your cluster to run this script when nodes are provisioned. Example script:
#!/bin/bash pip install spark-sklearn==0.2.0
Step 3: Verify the Installation
After installation, confirm that best_score_ is available by running a quick test in Python on any node:
from spark_sklearn.grid_search import GridSearchCV from sklearn.svm import SVC from sklearn.datasets import load_iris # Create a simple GridSearchCV instance iris = load_iris() param_grid = {'C': [1, 10]} gs = GridSearchCV(SVC(), param_grid) # Check if best_score_ exists print('best_score_ exists:', hasattr(gs, 'best_score_'))
If this returns True, you're good to go.
内容的提问来源于stack exchange,提问作者Falconic

