You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Python中Random Forest Regression的ROAS预测效率?

问题

我用随机森林回归模型,根据给定的TV、Radio、Newspaper广告成本上限计算ROAS(广告支出回报率),目标是找到能让模型输出Sales最高的成本组合。目前用三重循环逐美元遍历所有可能组合,30美元量级的输入就要跑5分钟,而且每次循环都重复加载模型,效率极低。

我的ROAS预测函数代码:

def ROASPrediction(Q,TV,Radio,Newspaper):
    rec = "Recommended Investment for Best Sales"
    y_big = 0
    x_b = 0
    y_b = 0
    z_b = 0
    for x in range(TV // 2, TV):
        for y in range(Radio // 2, Radio):
            for z in range(Newspaper // 2, Newspaper):
                customer_features = np.array([x, y, z])
                customer_features1 = customer_features.reshape(1, -1)
                # customer_features1 =pd.DataFrame(customer_features)
                model_fit1 = joblib.load('/content/drive/MyDrive/ROAS.joblib')
                y_future_pred = model_fit1.predict(customer_features1)
                print("y_future_pred", y_future_pred)
                if (y_future_pred[0] >= y_big):
                    y_big = y_future_pred[0]
                    x_b = x
                    y_b = y
                    z_b = z
    # y_future_pred1= str(y_future_pred[0]) + "M$"
    # y_roas= y_future_pred[0]*1000000 / (TV+Radio+Newspaper)
    y_future_pred1 = str(y_big) + "M$"
    y_roas = y_big * 1000000 / (TV + Radio + Newspaper)
    x_b1 = str(x_b)
    y_b1 = str(y_b)
    z_b1 = str(z_b)
    y_roas1 = str(y_roas) + "%"
    return rec, x_b1, y_b1, z_b1, y_future_pred1, y_roas1

模型训练代码:

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestRegressor

df = pd.read_csv('/Advertising.csv')
x = df[['TV', 'Radio','Newspaper']]
y = df[['Sales']]
x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.20, random_state=41)
rf_regressor = RandomForestRegressor(n_estimators=100, random_state=42)
rf_regressor.fit(x_train, y_train)
y_pred = rf_regressor.predict(x_test)

使用的是广告销售数据集,求更高效的方法提升运行速度。


优化方案

1. 紧急修复:移除重复加载模型逻辑

每次循环加载模型是最大性能瓶颈,将模型加载移到函数外部(全局或调用前执行),避免重复IO操作:

# 提前加载模型,仅执行一次
model_fit1 = joblib.load('/content/drive/MyDrive/ROAS.joblib')

def ROASPrediction(Q,TV,Radio,Newspaper):
    rec = "Recommended Investment for Best Sales"
    y_big = 0
    x_b = 0
    y_b = 0
    z_b = 0
    # 移除循环内的模型加载代码
    for x in range(TV // 2, TV):
        for y in range(Radio // 2, Radio):
            for z in range(Newspaper // 2, Newspaper):
                customer_features = np.array([x, y, z]).reshape(1, -1)
                y_future_pred = model_fit1.predict(customer_features)
                # 移除不必要的print输出,减少IO耗时
                if y_future_pred[0] >= y_big:
                    y_big = y_future_pred[0]
                    x_b, y_b, z_b = x, y, z
    # 后续格式化逻辑不变
    y_future_pred1 = f"{y_big}M$"
    y_roas = y_big * 1000000 / (TV + Radio + Newspaper)
    return (
        rec, str(x_b), str(y_b), str(z_b),
        y_future_pred1, f"{y_roas}%"
    )

2. 核心优化:批量生成特征+批量预测

用numpy向量运算替代三重循环,一次性生成所有特征组合并批量预测,速度提升几个数量级:

import numpy as np

model_fit1 = joblib.load('/content/drive/MyDrive/ROAS.joblib')

def ROASPrediction(Q,TV,Radio,Newspaper):
    rec = "Recommended Investment for Best Sales"
    # 生成所有可能的成本组合网格
    x_vals = np.arange(TV//2, TV)
    y_vals = np.arange(Radio//2, Radio)
    z_vals = np.arange(Newspaper//2, Newspaper)
    X_grid = np.array(np.meshgrid(x_vals, y_vals, z_vals)).T.reshape(-1, 3)
    # 批量预测所有组合
    predictions = model_fit1.predict(X_grid)
    # 找到最大值对应的组合
    max_idx = np.argmax(predictions)
    y_big = predictions[max_idx]
    x_b, y_b, z_b = X_grid[max_idx].astype(int)
    # 格式化输出
    y_future_pred1 = f"{y_big}M$"
    y_roas = y_big * 1000000 / (TV + Radio + Newspaper)
    return (
        rec, str(x_b), str(y_b), str(z_b),
        y_future_pred1, f"{y_roas}%"
    )

30美元量级的输入仅需几秒即可完成计算。

3. 进阶优化:用智能搜索替代全网格遍历

如果输入上限较大(如数百美元),全网格搜索仍会占用过多资源,可使用贝叶斯优化高效搜索最优解:

from bayes_opt import BayesianOptimization

model_fit1 = joblib.load('/content/drive/MyDrive/ROAS.joblib')

def sales_objective(x, y, z, TV, Radio, Newspaper):
    # 约束变量在指定范围内
    x = np.clip(int(x), TV//2, TV-1)
    y = np.clip(int(y), Radio//2, Radio-1)
    z = np.clip(int(z), Newspaper//2, Newspaper-1)
    return model_fit1.predict(np.array([[x, y, z]]))[0]

def ROASPrediction(Q,TV,Radio,Newspaper):
    rec = "Recommended Investment for Best Sales"
    # 定义搜索边界
    pbounds = {
        'x': (TV//2, TV-1),
        'y': (Radio//2, Radio-1),
        'z': (Newspaper//2, Newspaper-1)
    }
    # 初始化贝叶斯优化器
    optimizer = BayesianOptimization(
        f=lambda x,y,z: sales_objective(x,y,z,TV,Radio,Newspaper),
        pbounds=pbounds,
        random_state=42
    )
    # 运行优化(仅需几十次迭代)
    optimizer.maximize(init_points=5, n_iter=20)
    # 获取最优结果
    max_result = optimizer.max
    y_big = max_result['target']
    x_b, y_b, z_b = [int(v) for v in max_result['params'].values()]
    # 格式化输出
    y_future_pred1 = f"{y_big}M$"
    y_roas = y_big * 1000000 / (TV + Radio + Newspaper)
    return (
        rec, str(x_b), str(y_b), str(z_b),
        y_future_pred1, f"{y_roas}%"
    )

需先安装依赖库:pip install bayesian-optimization,该方法无需遍历所有组合,即可快速找到近似最优解。


内容的提问来源于stack exchange,提问作者l_b

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 23:54:56