如何整理多维金融时序与基本面数据至单DataFrame用于ANN股票预测清洗?
股票数据合并、清洗到ANN建模实操指南
一、将多股票数据合并为单个DataFrame
根据数据存储方式选择对应合并方案:
1. 数据自带股票标识列
若每只股票的数据源已包含stock_id(或股票代码)列,直接批量读取后拼接:
import pandas as pd import glob # 批量读取所有股票数据文件(假设为CSV格式) file_list = glob.glob("/path/to/your/stock_data/*.csv") df_list = [pd.read_csv(file) for file in file_list] # 合并为单个DataFrame combined_df = pd.concat(df_list, ignore_index=True)
2. 按文件名区分股票
若文件名本身是股票代码(如AAPL.csv、MSFT.csv),读取时手动添加stock_id列:
file_list = glob.glob("/path/to/your/stock_data/*.csv") df_list = [] for file_path in file_list: # 从文件名提取股票代码 stock_id = file_path.split("/")[-1].replace(".csv", "") temp_df = pd.read_csv(file_path) temp_df["stock_id"] = stock_id df_list.append(temp_df) combined_df = pd.concat(df_list, ignore_index=True)
3. 设置复合索引(推荐)
为方便后续按股票和时间分组处理,将stock_id和日期列设为复合索引:
# 转换日期列为datetime格式 combined_df["date"] = pd.to_datetime(combined_df["date"]) # 设置复合索引并排序 combined_df = combined_df.set_index(["stock_id", "date"]).sort_index()
二、数据清洗
针对股票时间序列数据的常见问题,按以下步骤处理:
1. 处理缺失值
股票数据缺失多因停牌、数据未收录导致,优先按股票分组做前后填充:
# 查看各特征缺失情况 print(combined_df.isnull().sum()) # 按股票分组,先前向填充再后向填充 combined_df = combined_df.groupby("stock_id").apply(lambda x: x.fillna(method="ffill").fillna(method="bfill")) # 若仍有剩余缺失值,直接删除对应行 combined_df = combined_df.dropna()
2. 剔除异常值
用百分位截断法处理极端值(避免干扰模型稳定性):
def truncate_outliers(df, col): # 取1%和99%分位数作为截断边界 lower = df[col].quantile(0.01) upper = df[col].quantile(0.99) return df[(df[col] >= lower) & (df[col] <= upper)] # 按股票分组处理每个特征的异常值 for feature in combined_df.columns: combined_df = combined_df.groupby("stock_id").apply(lambda x: truncate_outliers(x, feature)).reset_index(drop=True) # 重新设置复合索引 combined_df = combined_df.set_index(["stock_id", "date"]).sort_index()
3. 特征标准化
ANN对数据尺度敏感,按股票分组做标准化(避免不同股票价格区间差异干扰):
from sklearn.preprocessing import StandardScaler feature_cols = combined_df.columns.tolist() # 假设所有列都是特征(后续将添加目标列) scaler = StandardScaler() # 按股票分组标准化特征 combined_df[feature_cols] = combined_df.groupby("stock_id")[feature_cols].transform(lambda x: scaler.fit_transform(x))
三、特征选择(降维)
针对11个特征,推荐两种实用降维方法:
1. 相关性过滤
删除高度相关的特征(如开盘价和收盘价通常相关性极高):
# 计算特征相关矩阵 corr_matrix = combined_df.corr() # 找出相关系数绝对值>0.9的特征对 high_corr_pairs = [(i, j) for i in corr_matrix.columns for j in corr_matrix.columns if i < j and abs(corr_matrix[i][j]) > 0.9] print("高度相关特征对:", high_corr_pairs) # 示例:删除其中一个重复特征(如删除开盘价) combined_df = combined_df.drop("open_price", axis=1)
2. 基于模型的特征选择
用随机森林或SelectKBest筛选对预测目标贡献大的特征:
from sklearn.feature_selection import SelectKBest, f_regression from sklearn.ensemble import RandomForestRegressor # 创建预测目标:下一周的收盘价(时间序列移位) combined_df["target"] = combined_df.groupby("stock_id")["close_price"].shift(-1) # 删除最后一行(移位后目标值为空) combined_df = combined_df.dropna() X = combined_df.drop("target", axis=1) y = combined_df["target"] # 方法1:SelectKBest选择Top 8特征 selector = SelectKBest(f_regression, k=8) X_selected = selector.fit_transform(X, y) selected_features = X.columns[selector.get_support()].tolist() print("选中的特征:", selected_features) # 方法2:随机森林特征重要性排序 rf = RandomForestRegressor(n_estimators=100, random_state=42) rf.fit(X, y) importance_df = pd.DataFrame({ "feature": X.columns, "importance": rf.feature_importances_ }).sort_values(by="importance", ascending=False) print(importance_df) # 选择Top 8特征 top_features = importance_df["feature"].head(8).tolist() X_selected = X[top_features]
四、ANN模型构建
注意时间序列数据不能随机划分训练集/测试集,需按时间分割:
import numpy as np from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Dropout # 按时间分割:前18年训练,后2年测试(根据你的数据时间范围调整) split_date = pd.to_datetime("2022-01-01") train_data = combined_df[combined_df.index.get_level_values("date") < split_date] test_data = combined_df[combined_df.index.get_level_values("date") >= split_date] # 准备训练/测试数据 X_train = train_data[top_features].values y_train = train_data["target"].values X_test = test_data[top_features].values y_test = test_data["target"].values # 构建ANN回归模型 model = Sequential([ Dense(64, activation="relu", input_shape=(len(top_features),)), Dropout(0.2), # 防止过拟合 Dense(32, activation="relu"), Dropout(0.2), Dense(1) # 回归任务,输出预测收盘价 ]) model.compile(optimizer="adam", loss="mean_squared_error") # 训练模型 history = model.fit(X_train, y_train, epochs=50, batch_size=32, validation_split=0.1) # 评估模型 test_mse = model.evaluate(X_test, y_test) print(f"测试集MSE:{test_mse:.4f}")
内容的提问来源于stack exchange,提问作者Zouari Farah
相关产品推荐
相关产品推荐

