You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何整理多维金融时序与基本面数据至单DataFrame用于ANN股票预测清洗?

股票数据合并、清洗到ANN建模实操指南

一、将多股票数据合并为单个DataFrame

根据数据存储方式选择对应合并方案:

1. 数据自带股票标识列

若每只股票的数据源已包含stock_id(或股票代码)列,直接批量读取后拼接:

import pandas as pd
import glob

# 批量读取所有股票数据文件(假设为CSV格式)
file_list = glob.glob("/path/to/your/stock_data/*.csv")
df_list = [pd.read_csv(file) for file in file_list]

# 合并为单个DataFrame
combined_df = pd.concat(df_list, ignore_index=True)

2. 按文件名区分股票

若文件名本身是股票代码(如AAPL.csv、MSFT.csv),读取时手动添加stock_id列:

file_list = glob.glob("/path/to/your/stock_data/*.csv")
df_list = []

for file_path in file_list:
    # 从文件名提取股票代码
    stock_id = file_path.split("/")[-1].replace(".csv", "")
    temp_df = pd.read_csv(file_path)
    temp_df["stock_id"] = stock_id
    df_list.append(temp_df)

combined_df = pd.concat(df_list, ignore_index=True)

3. 设置复合索引(推荐)

为方便后续按股票和时间分组处理,将stock_id和日期列设为复合索引:

# 转换日期列为datetime格式
combined_df["date"] = pd.to_datetime(combined_df["date"])
# 设置复合索引并排序
combined_df = combined_df.set_index(["stock_id", "date"]).sort_index()

二、数据清洗

针对股票时间序列数据的常见问题,按以下步骤处理:

1. 处理缺失值

股票数据缺失多因停牌、数据未收录导致,优先按股票分组做前后填充:

# 查看各特征缺失情况
print(combined_df.isnull().sum())

# 按股票分组,先前向填充再后向填充
combined_df = combined_df.groupby("stock_id").apply(lambda x: x.fillna(method="ffill").fillna(method="bfill"))

# 若仍有剩余缺失值,直接删除对应行
combined_df = combined_df.dropna()

2. 剔除异常值

用百分位截断法处理极端值(避免干扰模型稳定性):

def truncate_outliers(df, col):
    # 取1%和99%分位数作为截断边界
    lower = df[col].quantile(0.01)
    upper = df[col].quantile(0.99)
    return df[(df[col] >= lower) & (df[col] <= upper)]

# 按股票分组处理每个特征的异常值
for feature in combined_df.columns:
    combined_df = combined_df.groupby("stock_id").apply(lambda x: truncate_outliers(x, feature)).reset_index(drop=True)

# 重新设置复合索引
combined_df = combined_df.set_index(["stock_id", "date"]).sort_index()

3. 特征标准化

ANN对数据尺度敏感,按股票分组做标准化(避免不同股票价格区间差异干扰):

from sklearn.preprocessing import StandardScaler

feature_cols = combined_df.columns.tolist()  # 假设所有列都是特征(后续将添加目标列)
scaler = StandardScaler()

# 按股票分组标准化特征
combined_df[feature_cols] = combined_df.groupby("stock_id")[feature_cols].transform(lambda x: scaler.fit_transform(x))

三、特征选择(降维)

针对11个特征,推荐两种实用降维方法:

1. 相关性过滤

删除高度相关的特征(如开盘价和收盘价通常相关性极高):

# 计算特征相关矩阵
corr_matrix = combined_df.corr()

# 找出相关系数绝对值>0.9的特征对
high_corr_pairs = [(i, j) for i in corr_matrix.columns for j in corr_matrix.columns 
                   if i < j and abs(corr_matrix[i][j]) > 0.9]
print("高度相关特征对:", high_corr_pairs)

# 示例:删除其中一个重复特征(如删除开盘价)
combined_df = combined_df.drop("open_price", axis=1)

2. 基于模型的特征选择

用随机森林或SelectKBest筛选对预测目标贡献大的特征:

from sklearn.feature_selection import SelectKBest, f_regression
from sklearn.ensemble import RandomForestRegressor

# 创建预测目标:下一周的收盘价(时间序列移位)
combined_df["target"] = combined_df.groupby("stock_id")["close_price"].shift(-1)
# 删除最后一行(移位后目标值为空)
combined_df = combined_df.dropna()

X = combined_df.drop("target", axis=1)
y = combined_df["target"]

# 方法1:SelectKBest选择Top 8特征
selector = SelectKBest(f_regression, k=8)
X_selected = selector.fit_transform(X, y)
selected_features = X.columns[selector.get_support()].tolist()
print("选中的特征:", selected_features)

# 方法2:随机森林特征重要性排序
rf = RandomForestRegressor(n_estimators=100, random_state=42)
rf.fit(X, y)
importance_df = pd.DataFrame({
    "feature": X.columns,
    "importance": rf.feature_importances_
}).sort_values(by="importance", ascending=False)
print(importance_df)

# 选择Top 8特征
top_features = importance_df["feature"].head(8).tolist()
X_selected = X[top_features]

四、ANN模型构建

注意时间序列数据不能随机划分训练集/测试集,需按时间分割:

import numpy as np
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Dropout

# 按时间分割:前18年训练,后2年测试(根据你的数据时间范围调整)
split_date = pd.to_datetime("2022-01-01")
train_data = combined_df[combined_df.index.get_level_values("date") < split_date]
test_data = combined_df[combined_df.index.get_level_values("date") >= split_date]

# 准备训练/测试数据
X_train = train_data[top_features].values
y_train = train_data["target"].values
X_test = test_data[top_features].values
y_test = test_data["target"].values

# 构建ANN回归模型
model = Sequential([
    Dense(64, activation="relu", input_shape=(len(top_features),)),
    Dropout(0.2),  # 防止过拟合
    Dense(32, activation="relu"),
    Dropout(0.2),
    Dense(1)  # 回归任务,输出预测收盘价
])

model.compile(optimizer="adam", loss="mean_squared_error")
# 训练模型
history = model.fit(X_train, y_train, epochs=50, batch_size=32, validation_split=0.1)

# 评估模型
test_mse = model.evaluate(X_test, y_test)
print(f"测试集MSE:{test_mse:.4f}")

内容的提问来源于stack exchange,提问作者Zouari Farah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 02:54:22