如何在按sku分组的回归分析中集成true.revenue与true.profit函数?
分组回归与自定义收益/利润函数整合方案
问题背景
给定包含sku、价格、需求、成本等字段的数据集,需要按sku分组拟合hist.demand ~ hist.prices的回归模型,计算最优收益价格和最优利润价格,并使用自定义函数true.revenue和true.profit计算对应最优值,但原代码因函数依赖未明确传入参数而报错,需调整流程实现整合。
解决方案
1. 修正自定义函数定义
原函数依赖未显式传入的全局变量,需修改为接收回归系数、单位成本等参数的形式,确保在分组计算时能正确获取每个组的对应值:
# 计算最优收益:p * 需求,需求由回归模型demand = intercept + slope*p得到 true.revenue <- function(p, intercept, slope) { demand <- intercept + slope * p p * demand } # 计算最优利润:(p - 单位成本) * 需求 true.profit <- function(p, intercept, slope, unity_cost) { demand <- intercept + slope * p (p - unity_cost) * demand }
2. 调整分组回归流程
使用dplyr的group_by() + reframe()语法(替代原nest操作),按sku分组拟合模型、提取系数、计算最优价格及对应收益/利润:
library(dplyr) library(tidyr) # 加载数据集 dat <- structure(list(sku = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L), period = c("30.09.2021", "14.03.2019", "01.04.2022", "18.02.2022", "07.07.2021", "09.10.2020", "17.01.2019", "10.11.2020", "14.07.2021", "10.09.2019", "31.01.2019", "01.07.2021", "30.09.2021", "14.03.2019", "01.04.2022", "18.02.2022", "07.07.2021", "09.10.2020", "17.01.2019", "10.11.2020", "14.07.2021", "10.09.2019", "31.01.2019", "01.07.2021"), hist.prices = c(3728.16, 34899.84, 6126, 1789.44, 18098.4, 15633.6, 26174.88, 2401.56, 12668.88, 239500.8, 26174.88, 5429.52, 3728.16, 34899.84, 6126, 1789.44, 18098.4, 15633.6, 26174.88, 2401.56, 12668.88, 239500.8, 26174.88, 5429.52), hist.revenue = c(178951.68, 20102307.84, 367560, 42946.56, 4343616, 3752064, 11307548.16, 86456.16, 2128371.84, 965667225.6, 11307548.16, 390925.44, 178951.68, 20102307.84, 367560, 42946.56, 4343616, 3752064, 11307548.16, 86456.16, 2128371.84, 965667225.6, 11307548.16, 390925.44), hist.demand = c(254L, 276L, 272L, 250L, 299L, 297L, 291L, 260L, 270L, 275L, 295L, 279L, 254L, 276L, 272L, 250L, 299L, 297L, 291L, 260L, 270L, 275L, 295L, 279L ), hist.cost = c(12572.6698, 10498.9848, 14949.392, 13160.5, 14557.9512, 12443.3199, 10692.3294, 10893.116, 13145.976, 10222.6025, 10982.9975, 13584.1752, 12572.6698, 10498.9848, 14949.392, 13160.5, 14557.9512, 12443.3199, 10692.3294, 10893.116, 13145.976, 10222.6025, 10982.9975, 13584.1752), unity.cost = c(49.4987, 38.0398, 54.961, 52.642, 48.6888, 41.8967, 36.7434, 41.8966, 48.6888, 37.1731, 37.2305, 48.6888, 49.4987, 38.0398, 54.961, 52.642, 48.6888, 41.8967, 36.7434, 41.8966, 48.6888, 37.1731, 37.2305, 48.6888 ), hist.profit = c(1336L, 1592L, 1128L, 1882L, 1387L, 1818L, 1357L, 1087L, 1253L, 1009L, 1092L, 1804L, 1336L, 1592L, 1128L, 1882L, 1387L, 1818L, 1357L, 1087L, 1253L, 1009L, 1092L, 1804L )), class = "data.frame", row.names = c(NA, -24L)) # 执行分组计算 result <- dat %>% group_by(sku) %>% reframe( # 拟合每个sku的需求-价格回归模型 model = list(lm(hist.demand ~ hist.prices)), # 提取回归截距和斜率(对应你描述的x和beta系数) intercept = coef(model[[1]])[["(Intercept)"]], slope = coef(model[[1]])[["hist.prices"]], # 处理单位成本:取该sku的平均单位成本(可根据需求调整为最新值/中位数等) avg_unity_cost = mean(unity.cost), # 计算最优收益价格和最优利润价格的公式 p.revenue = -intercept / (2 * slope), p.profit = (slope * avg_unity_cost - intercept) / (2 * slope), # 调用自定义函数计算最优收益和利润 opt.revenue = true.revenue(p.revenue, intercept, slope), opt.profit = true.profit(p.profit, intercept, slope, avg_unity_cost) ) %>% # 移除模型列(可选,根据需要保留) select(-model) # 查看结果 print(result)
3. (可选)将最优值合并回原数据集
如果需要每个原始数据行都带上对应sku的最优值,可使用left_join:
dat_with_opt <- dat %>% left_join(result, by = "sku")
关键点说明
- 参数显式传递:自定义函数必须接收
intercept(回归截距)、slope(回归斜率)、unity_cost(单位成本)等参数,避免依赖全局变量,确保分组计算时每个组使用自身的系数。 - 单位成本处理:你的数据中每个sku的
unity.cost存在多行不同值,示例中取平均值,可根据业务需求调整为最新记录值、中位数或其他统计量。 - 回归系数命名:原代码中
beta和alpha的命名易混淆,示例中明确用intercept(截距)和slope(斜率)对应你描述的x和beta系数,提升代码可读性。
内容的提问来源于stack exchange,提问作者psysky
相关产品推荐
相关产品推荐

