在Jupyter Lab运行Scikit-learn加州房价预测代码时遭遇SSL证书验证失败错误的求助
Hey there, let's sort out this SSL certificate error you're running into while trying to load the California Housing dataset. The root cause here is that your Python environment can't verify the SSL certificate for the dataset's download server—this is a super common issue, especially on macOS, or if you're working behind a corporate proxy.
First, let's break down the error you received:
URLError: <urlopen error [SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: unable to get local issuer certificate (_ssl.c:1077)>
This means the SSL handshake between your machine and the dataset server failed because your system doesn't trust the organization that issued the server's security certificate.
Easy Fixes to Try
1. Temporarily Disable SSL Verification
You can skip the SSL check just for this dataset fetch by adding the ssl_verify=False parameter to the fetch_california_housing call. Also, I noticed a tiny typo in your code—you tried to print rms but your variable is named rmse; I fixed that in the snippet below:
from sklearn.datasets import fetch_california_housing from sklearn.model_selection import train_test_split from sklearn.ensemble import RandomForestRegressor from sklearn.metrics import mean_squared_error import numpy as np # Add ssl_verify=False to bypass SSL certificate check california_housing = fetch_california_housing(ssl_verify=False) X = california_housing.data y = california_housing.target X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42) rf_reg = RandomForestRegressor(random_state=42) rf_reg.fit(X_train, y_train) y_pred = rf_reg.predict(X_test) mse = mean_squared_error(y_test, y_pred) rmse = np.sqrt(mse) # Fixed the typo here: print(rmse) instead of print(rms) print(rmse)
2. Update Your System's SSL Certificates (Safer Long-Term)
If you don't want to disable SSL verification (which is the more secure approach), update your system's root certificates:
- macOS: Find the
Install Certificates.commandfile in your Python installation folder (usually/Applications/Python X.X/where X.X is your Python version). Double-click it to run the script, which installs the necessary root certificates. - Windows: Upgrade the
certifipackage (which handles SSL certificates in Python) using pip:pip install --upgrade certifi - Linux: Use your system's package manager to refresh the CA certificates. For Ubuntu/Debian:
sudo apt-get update && sudo apt-get install --reinstall ca-certificates
3. Download the Dataset Manually
If the above options don't work, you can grab the dataset file directly and load it locally:
- Download the
california_housing.tgzfile from the scikit-learn dataset repository. - Extract the file to get
california_housing.npz. - Load the dataset using
np.loadwith your local file path:import numpy as np from sklearn.model_selection import train_test_split from sklearn.ensemble import RandomForestRegressor from sklearn.metrics import mean_squared_error # Replace with the actual path to your extracted .npz file data = np.load('/your/path/to/california_housing.npz') X = data['data'] y = data['target'] # Rest of your training code stays the same X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42) rf_reg = RandomForestRegressor(random_state=42) rf_reg.fit(X_train, y_train) y_pred = rf_reg.predict(X_test) mse = mean_squared_error(y_test, y_pred) rmse = np.sqrt(mse) print(rmse)
内容的提问来源于stack exchange,提问作者omersimsir

