Pandas IndexError报错求助:使用Matplotlib绘制能源负荷实际值与预测值时的索引问题
Let's break down what's causing this error and how to fix it quickly.
The Root Cause
The error message spells it out clearly: .iloc doesn't accept a DataFrame as an index. When you run train_test_split, the returned X_test is a DataFrame that retains the original row indices from your data dataset. Trying to use data['timestamp'].iloc[X_test] fails because .iloc expects integer positions (like 0, 1, 2), not an entire DataFrame.
The Fix
Instead of passing the whole X_test DataFrame to .iloc, you need to use the row index labels from X_test to fetch the corresponding timestamps from your original data. Here's how to adjust your plotting code:
plt.figure(figsize=(15, 5)) # Use X_test's index to get matching timestamps from the original data plt.scatter(data['timestamp'].loc[X_test.index], y_test, s=5, label="Actual") plt.scatter(data['timestamp'].loc[X_test.index], y_pred, s=5, color='red', label="Predicted") plt.xlabel("Datetime") plt.ylabel("Load") plt.title("Energy Load:Actual vs Predicted") plt.legend() plt.tight_layout() plt.show()
Why This Works
X_test.indexgives you the exact row/timestamp labels from your originaldatadataset that were assigned to the test set..locis designed to access data using these index labels, which aligns perfectly with what the error message suggested ("Consider using .loc for automatic alignment").
What Might Have Gone Wrong With Your Earlier Attempt
If you tried converting X_test to a list directly, you were probably getting the column names (like ['time_num', 'temp']) instead of the row indices. Always use the .index attribute when you need the row labels of a DataFrame.
Full Modified Code
For completeness, here's your entire code with the fix applied:
import pandas as pd import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression path = '/content/drive/MyDrive/ML_DATASETS/energy.csv' data = pd.read_csv(path) data['timestamp'] = pd.to_datetime(data['timestamp']) data['time_num'] = range(len(data)) X = data[['time_num', 'temp']] y = data['load'] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2) model = LinearRegression() model.fit(X_train, y_train) y_pred = model.predict(X_test) # Fixed plotting code plt.figure(figsize=(15, 5)) plt.scatter(data['timestamp'].loc[X_test.index], y_test, s=5, label="Actual") plt.scatter(data['timestamp'].loc[X_test.index], y_pred, s=5, color='red', label="Predicted") plt.xlabel("Datetime") plt.ylabel("Load") plt.title("Energy Load:Actual vs Predicted") plt.legend() plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者Rima 123

