如何在Python中分批迭代两个DataFrame行进行回归计算?
分批次处理训练/验证集的示例代码
直接用iloc切片实现批量划分是最简洁高效的方式,避免iterrows逐行遍历的低效问题。以下是完整示例:
示例代码
import pandas as pd import numpy as np from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error # ---------------------- # 1. 模拟生成训练/验证数据(替换成你自己的加载代码) # ---------------------- # 假设特征列是x1-x5,目标列是y np.random.seed(42) train_df = pd.DataFrame( np.random.rand(40000, 6), columns=[f'x{i}' for i in range(1,6)] + ['y'] ) val_df = pd.DataFrame( np.random.rand(40000, 6), columns=[f'x{i}' for i in range(1,6)] + ['y'] ) # ---------------------- # 2. 批量处理逻辑 # ---------------------- batch_size = 1000 total_batches = 40 error_history = [] for batch_idx in range(total_batches): # 计算当前批次的起始/结束索引 start = batch_idx * batch_size end = start + batch_size # 提取当前批次的训练/验证数据 train_batch = train_df.iloc[start:end] val_batch = val_df.iloc[start:end] # 分离特征(X)和目标(y) X_train = train_batch.drop('y', axis=1) y_train = train_batch['y'] X_val = val_batch.drop('y', axis=1) y_val = val_batch['y'] # 训练回归模型(这里用线性回归,替换成你自己的模型) model = LinearRegression() model.fit(X_train, y_train) # 预测并计算误差 y_pred = model.predict(X_val) mse = mean_squared_error(y_val, y_pred) error_history.append(mse) # 打印当前批次结果 print(f"批次 {batch_idx+1}/{total_batches} - MSE: {mse:.4f}") # 所有批次处理完成后,查看误差历史 print("\n所有批次误差列表:", error_history)
关键说明
- 批量划分:用
iloc[start:end]直接截取连续的1000行,比iterrows高效得多,且逻辑清晰。 - 模型训练:示例中每次批次都重新初始化模型,如果你需要增量训练(用之前批次的模型继续训练当前批次),只需把
model = LinearRegression()移到循环外面即可。 - 误差计算:用
mean_squared_error作为示例,你可以替换成MAE、RMSE等其他误差指标。
内容的提问来源于stack exchange,提问作者BlueSky01
相关产品推荐
相关产品推荐

