Python模型回测时触发ValueError: No objects to concatenate错误求助
解决回测代码中的「ValueError: No objects to concatenate」错误
问题概述
使用MSFT股价数据回测随机森林分类模型时,运行代码触发ValueError: No objects to concatenate错误,预期输出回测中预测价格上涨(标记1)和下跌(标记0)的数量统计。
完整出错代码
import yfinance as yf import numpy as np import pandas as pd sp_500 = yf.Ticker("MSFT") sp_500 = sp_500.history(period="max") # sp_500.plot.line(y="Close", use_index=True) del sp_500["Dividends"] del sp_500["Stock Splits"] sp_500["Tomorrow"] = sp_500["Close"].shift(-1) sp_500["Target"] = (sp_500["Tomorrow"] > sp_500["Close"]).astype(int) # 转换为0/1而非布尔值 sp_500 = sp_500.loc["2022-01-01":].copy() from sklearn.ensemble import RandomForestClassifier from sklearn.metrics import precision_score model = RandomForestClassifier(n_estimators=100, min_samples_split=100, random_state=1) train = sp_500.iloc[:-100] test = sp_500.iloc[-100:] predictors = ["Close", "Volume", "Open", "High", "Low"] model.fit(train[predictors], train["Target"]) preds = model.predict(test[predictors]) preds = pd.Series(preds, index=test.index) combined = pd.concat([test["Target"], preds], axis=1) def predict(train, test, predictors, model): model.fit(train[predictors], train["Target"]) preds = model.predict(test[predictors]) preds = pd.Series(preds, index=test.index, name="Predictions") combined = pd.concat([test["Target"], preds], axis=1) return combined def backtest(data, model, predictors, start=2500, step=250): all_predictions = [] for i in range(start, data.shape[0] - step, step): train = data.iloc[i:(i + step)] test = data.iloc[(i + step):(i + 2 * step)] predictions = predict(train, test, predictors, model) all_predictions.append(predictions) return pd.concat(all_predictions) predictions = backtest(sp_500, model, predictors) print(predictions["Predictions"].value_counts())
错误原因
核心问题是回测参数与现有数据量不匹配:
- 代码筛选了2022年至今的MSFT数据,这段时间的交易日数量仅约600行左右。
- 但
backtest函数默认设置start=2500,远大于现有数据的总行数,导致range(start, data.shape[0]-step, step)生成的循环范围为空,all_predictions始终是一个空列表。 - 调用
pd.concat(all_predictions)时,传入空列表就会触发No objects to concatenate错误。
解决方法
步骤1:确认现有数据量
先添加代码查看筛选后的数据行数,明确数据规模:
print(sp_500.shape[0])
步骤2:调整回测参数
修改backtest函数的start和step参数,确保循环能正常执行。比如将start设为200,step设为50(数值可根据实际数据量灵活调整):
方式一:修改函数默认参数
def backtest(data, model, predictors, start=200, step=50): all_predictions = [] for i in range(start, data.shape[0] - step, step): train = data.iloc[i:(i + step)] test = data.iloc[(i + step):(i + 2 * step)] predictions = predict(train, test, predictors, model) all_predictions.append(predictions) return pd.concat(all_predictions)
方式二:调用时指定参数
predictions = backtest(sp_500, model, predictors, start=200, step=50)
步骤3:验证修改效果
运行修改后的代码,即可正常生成回测预测结果,并输出Predictions列中1和0的计数。
内容的提问来源于stack exchange,提问作者Devansh
相关产品推荐
相关产品推荐

