如何在Python中为多品类多门店复用Ridge回归模型实现销售预测
Let's walk through a step-by-step approach to build and apply your Ridge regression model to predict weekly sales across all 4 depots and 4 products. We'll focus on time-series best practices (like avoiding random train/test splits) and systematic handling of each depot-product pair.
1. Setup & Data Preprocessing
First, we'll load and clean the data, then aggregate daily sales into weekly totals since your goal is weekly prediction.
import pandas as pd import numpy as np from sklearn.linear_model import Ridge from sklearn.metrics import mean_squared_error, r2_score # Load your dataset (replace with your file path) df = pd.read_csv("sales_data.csv") df["Date"] = pd.to_datetime(df["Date"]) # Aggregate daily sales to weekly totals df["Year"] = df["Date"].dt.year df["Week"] = df["Date"].dt.isocalendar().week weekly_sales = df.groupby(["DepotName", "Product", "Year", "Week"])["SalesUnits"].sum().reset_index() # Add a week start date for clarity weekly_sales["WeekStart"] = pd.to_datetime( weekly_sales["Year"].astype(str) + "-W" + weekly_sales["Week"].astype(str) + "-1", format="%Y-W%W-%w" )
2. Feature Engineering for Time Series
For linear models to work well on time-series data, we need relevant features. A simple but effective feature is lagged sales (previous week's sales), which captures trends. We'll also include year and week to account for seasonality.
# Create lag feature (previous week's sales) for each depot-product pair weekly_sales["Lag1"] = weekly_sales.groupby(["DepotName", "Product"])["SalesUnits"].shift(1) # Drop rows with missing lag values (first week of data has no prior sales) weekly_sales = weekly_sales.dropna(subset=["Lag1"])
3. Model Training: Two Approaches
You have two main options to apply the Ridge model to all combinations:
Option 1: Separate Model for Each Depot-Product Pair
This approach lets each combination have its own model, which can capture unique sales patterns for that specific pair.
# Store models and evaluation results models = {} evaluation_results = [] # Iterate over every unique depot-product combination for (depot, product), group in weekly_sales.groupby(["DepotName", "Product"]): # Split data into train (80% historical) and test (20% recent) - time-based split! train_size = int(0.8 * len(group)) X_train = group[["Lag1", "Year", "Week"]].iloc[:train_size] y_train = group["SalesUnits"].iloc[:train_size] X_test = group[["Lag1", "Year", "Week"]].iloc[train_size:] y_test = group["SalesUnits"].iloc[train_size:] # Train Ridge model reg = Ridge(alpha=1) reg.fit(X_train, y_train) # Evaluate performance y_pred = reg.predict(X_test) mse = mean_squared_error(y_test, y_pred) r2 = r2_score(y_test, y_pred) # Save model and results models[(depot, product)] = reg evaluation_results.append({ "Depot": depot, "Product": product, "TestMSE": round(mse, 2), "R2Score": round(r2, 2) }) # Print evaluation summary print(pd.DataFrame(evaluation_results))
Option 2: Single Model with Categorical Encoding
If you prefer a single model that handles all combinations, encode the depot and product as categorical features using one-hot encoding.
# One-hot encode categorical variables encoded_data = pd.get_dummies(weekly_sales, columns=["DepotName", "Product"], drop_first=True) # Split data (time-based split) train_size = int(0.8 * len(encoded_data)) X_train = encoded_data.drop(["SalesUnits", "WeekStart"], axis=1).iloc[:train_size] y_train = encoded_data["SalesUnits"].iloc[:train_size] X_test = encoded_data.drop(["SalesUnits", "WeekStart"], axis=1).iloc[train_size:] y_test = encoded_data["SalesUnits"].iloc[train_size:] # Train and evaluate single Ridge model reg = Ridge(alpha=1) reg.fit(X_train, y_train) y_pred = reg.predict(X_test) print(f"Overall Test MSE: {round(mean_squared_error(y_test, y_pred), 2)}") print(f"Overall R2 Score: {round(r2_score(y_test, y_pred), 2)}")
4. Predict Future Weekly Sales
Once your models are trained, you can predict sales for upcoming weeks. Here's how to do it with the separate models approach:
# Predict sales for the next week next_year = weekly_sales["Year"].max() next_week = weekly_sales["Week"].max() + 1 predictions = [] for (depot, product), model in models.items(): # Get the latest sales value to use as lag feature latest_sales = weekly_sales[ (weekly_sales["DepotName"] == depot) & (weekly_sales["Product"] == product) ]["SalesUnits"].iloc[-1] # Create feature data for next week X_pred = pd.DataFrame({ "Lag1": [latest_sales], "Year": [next_year], "Week": [next_week] }) # Generate prediction predicted_sales = model.predict(X_pred)[0] predictions.append({ "DepotName": depot, "Product": product, "Year": next_year, "Week": next_week, "PredictedSalesUnits": round(predicted_sales, 2) }) # Print predictions print(pd.DataFrame(predictions))
Key Notes
- Time-based splits: Never use random train/test splits for time-series data—this leaks future information into your training set.
- Feature expansion: You can add more features like month, holidays, or longer lag periods (e.g., Lag2 for sales two weeks prior) to improve model performance.
- Hyperparameter tuning: Consider using
GridSearchCVto optimize thealphaparameter for Ridge regression instead of hardcoding it to 1.
内容的提问来源于stack exchange,提问作者Ahamed Moosa

