You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas IndexError报错求助:使用Matplotlib绘制能源负荷实际值与预测值时的索引问题

Fixing IndexError When Plotting Actual vs Predicted Energy Load

Let's break down what's causing this error and how to fix it quickly.

The Root Cause

The error message spells it out clearly: .iloc doesn't accept a DataFrame as an index. When you run train_test_split, the returned X_test is a DataFrame that retains the original row indices from your data dataset. Trying to use data['timestamp'].iloc[X_test] fails because .iloc expects integer positions (like 0, 1, 2), not an entire DataFrame.

The Fix

Instead of passing the whole X_test DataFrame to .iloc, you need to use the row index labels from X_test to fetch the corresponding timestamps from your original data. Here's how to adjust your plotting code:

plt.figure(figsize=(15, 5))
# Use X_test's index to get matching timestamps from the original data
plt.scatter(data['timestamp'].loc[X_test.index], y_test, s=5, label="Actual")
plt.scatter(data['timestamp'].loc[X_test.index], y_pred, s=5, color='red', label="Predicted")
plt.xlabel("Datetime")
plt.ylabel("Load")
plt.title("Energy Load:Actual vs Predicted")
plt.legend()
plt.tight_layout()
plt.show()

Why This Works

  • X_test.index gives you the exact row/timestamp labels from your original data dataset that were assigned to the test set.
  • .loc is designed to access data using these index labels, which aligns perfectly with what the error message suggested ("Consider using .loc for automatic alignment").

What Might Have Gone Wrong With Your Earlier Attempt

If you tried converting X_test to a list directly, you were probably getting the column names (like ['time_num', 'temp']) instead of the row indices. Always use the .index attribute when you need the row labels of a DataFrame.

Full Modified Code

For completeness, here's your entire code with the fix applied:

import pandas as pd
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression

path = '/content/drive/MyDrive/ML_DATASETS/energy.csv'
data = pd.read_csv(path)
data['timestamp'] = pd.to_datetime(data['timestamp'])
data['time_num'] = range(len(data))

X = data[['time_num', 'temp']]
y = data['load']

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)

model = LinearRegression()
model.fit(X_train, y_train)
y_pred = model.predict(X_test)

# Fixed plotting code
plt.figure(figsize=(15, 5))
plt.scatter(data['timestamp'].loc[X_test.index], y_test, s=5, label="Actual")
plt.scatter(data['timestamp'].loc[X_test.index], y_pred, s=5, color='red', label="Predicted")
plt.xlabel("Datetime")
plt.ylabel("Load")
plt.title("Energy Load:Actual vs Predicted")
plt.legend()
plt.tight_layout()
plt.show()

内容的提问来源于stack exchange,提问作者Rima 123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 09:47:34