如何优化Python中Random Forest Regression的ROAS预测效率?
问题
我用随机森林回归模型,根据给定的TV、Radio、Newspaper广告成本上限计算ROAS(广告支出回报率),目标是找到能让模型输出Sales最高的成本组合。目前用三重循环逐美元遍历所有可能组合,30美元量级的输入就要跑5分钟,而且每次循环都重复加载模型,效率极低。
我的ROAS预测函数代码:
def ROASPrediction(Q,TV,Radio,Newspaper): rec = "Recommended Investment for Best Sales" y_big = 0 x_b = 0 y_b = 0 z_b = 0 for x in range(TV // 2, TV): for y in range(Radio // 2, Radio): for z in range(Newspaper // 2, Newspaper): customer_features = np.array([x, y, z]) customer_features1 = customer_features.reshape(1, -1) # customer_features1 =pd.DataFrame(customer_features) model_fit1 = joblib.load('/content/drive/MyDrive/ROAS.joblib') y_future_pred = model_fit1.predict(customer_features1) print("y_future_pred", y_future_pred) if (y_future_pred[0] >= y_big): y_big = y_future_pred[0] x_b = x y_b = y z_b = z # y_future_pred1= str(y_future_pred[0]) + "M$" # y_roas= y_future_pred[0]*1000000 / (TV+Radio+Newspaper) y_future_pred1 = str(y_big) + "M$" y_roas = y_big * 1000000 / (TV + Radio + Newspaper) x_b1 = str(x_b) y_b1 = str(y_b) z_b1 = str(z_b) y_roas1 = str(y_roas) + "%" return rec, x_b1, y_b1, z_b1, y_future_pred1, y_roas1
模型训练代码:
import pandas as pd from sklearn.model_selection import train_test_split from sklearn.ensemble import RandomForestRegressor df = pd.read_csv('/Advertising.csv') x = df[['TV', 'Radio','Newspaper']] y = df[['Sales']] x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.20, random_state=41) rf_regressor = RandomForestRegressor(n_estimators=100, random_state=42) rf_regressor.fit(x_train, y_train) y_pred = rf_regressor.predict(x_test)
使用的是广告销售数据集,求更高效的方法提升运行速度。
优化方案
1. 紧急修复:移除重复加载模型逻辑
每次循环加载模型是最大性能瓶颈,将模型加载移到函数外部(全局或调用前执行),避免重复IO操作:
# 提前加载模型,仅执行一次 model_fit1 = joblib.load('/content/drive/MyDrive/ROAS.joblib') def ROASPrediction(Q,TV,Radio,Newspaper): rec = "Recommended Investment for Best Sales" y_big = 0 x_b = 0 y_b = 0 z_b = 0 # 移除循环内的模型加载代码 for x in range(TV // 2, TV): for y in range(Radio // 2, Radio): for z in range(Newspaper // 2, Newspaper): customer_features = np.array([x, y, z]).reshape(1, -1) y_future_pred = model_fit1.predict(customer_features) # 移除不必要的print输出,减少IO耗时 if y_future_pred[0] >= y_big: y_big = y_future_pred[0] x_b, y_b, z_b = x, y, z # 后续格式化逻辑不变 y_future_pred1 = f"{y_big}M$" y_roas = y_big * 1000000 / (TV + Radio + Newspaper) return ( rec, str(x_b), str(y_b), str(z_b), y_future_pred1, f"{y_roas}%" )
2. 核心优化:批量生成特征+批量预测
用numpy向量运算替代三重循环,一次性生成所有特征组合并批量预测,速度提升几个数量级:
import numpy as np model_fit1 = joblib.load('/content/drive/MyDrive/ROAS.joblib') def ROASPrediction(Q,TV,Radio,Newspaper): rec = "Recommended Investment for Best Sales" # 生成所有可能的成本组合网格 x_vals = np.arange(TV//2, TV) y_vals = np.arange(Radio//2, Radio) z_vals = np.arange(Newspaper//2, Newspaper) X_grid = np.array(np.meshgrid(x_vals, y_vals, z_vals)).T.reshape(-1, 3) # 批量预测所有组合 predictions = model_fit1.predict(X_grid) # 找到最大值对应的组合 max_idx = np.argmax(predictions) y_big = predictions[max_idx] x_b, y_b, z_b = X_grid[max_idx].astype(int) # 格式化输出 y_future_pred1 = f"{y_big}M$" y_roas = y_big * 1000000 / (TV + Radio + Newspaper) return ( rec, str(x_b), str(y_b), str(z_b), y_future_pred1, f"{y_roas}%" )
30美元量级的输入仅需几秒即可完成计算。
3. 进阶优化:用智能搜索替代全网格遍历
如果输入上限较大(如数百美元),全网格搜索仍会占用过多资源,可使用贝叶斯优化高效搜索最优解:
from bayes_opt import BayesianOptimization model_fit1 = joblib.load('/content/drive/MyDrive/ROAS.joblib') def sales_objective(x, y, z, TV, Radio, Newspaper): # 约束变量在指定范围内 x = np.clip(int(x), TV//2, TV-1) y = np.clip(int(y), Radio//2, Radio-1) z = np.clip(int(z), Newspaper//2, Newspaper-1) return model_fit1.predict(np.array([[x, y, z]]))[0] def ROASPrediction(Q,TV,Radio,Newspaper): rec = "Recommended Investment for Best Sales" # 定义搜索边界 pbounds = { 'x': (TV//2, TV-1), 'y': (Radio//2, Radio-1), 'z': (Newspaper//2, Newspaper-1) } # 初始化贝叶斯优化器 optimizer = BayesianOptimization( f=lambda x,y,z: sales_objective(x,y,z,TV,Radio,Newspaper), pbounds=pbounds, random_state=42 ) # 运行优化(仅需几十次迭代) optimizer.maximize(init_points=5, n_iter=20) # 获取最优结果 max_result = optimizer.max y_big = max_result['target'] x_b, y_b, z_b = [int(v) for v in max_result['params'].values()] # 格式化输出 y_future_pred1 = f"{y_big}M$" y_roas = y_big * 1000000 / (TV + Radio + Newspaper) return ( rec, str(x_b), str(y_b), str(z_b), y_future_pred1, f"{y_roas}%" )
需先安装依赖库:pip install bayesian-optimization,该方法无需遍历所有组合,即可快速找到近似最优解。
内容的提问来源于stack exchange,提问作者l_b
相关产品推荐
相关产品推荐

