如何计算R dataframe中每个连续折扣组首尾行的价格差值
实现思路
核心逻辑是先为连续取值相同的discount序列生成唯一分组ID,再按组计算首尾价格的差值即可,下面提供两种常用实现方案:
方案1:tidyverse 实现(最简洁)
dplyr 1.1.0及以上版本自带consecutive_id()函数,专门用来给连续相同的取值生成分组ID,代码易读性很高:
library(tidyverse) # 你的示例数据 df <- structure(list(date = structure(c(16437, 16440, 16441, 16442,16443, 16444, 16447, 16448, 16449, 16450, 16451, 16454, 16455, 16456, 16457, 16458, 16461, 16462, 16463, 16464), class = "Date"), price = c(7.859, 7.964, 7.942, 8.117, 8.066, 7.918, 7.856, 7.82, 7.828, 7.791, 7.785, 7.789, 7.446, 7.401, 7.37, 7.283, 7.299, 7.25, 7.228, 7.219), discount = c(1, 1, 1, 1, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0)), row.names = c(NA, -20L), class = c("tbl_df", "tbl", "data.frame")) # 计算每个连续折扣组的首尾价差 result <- df %>% group_by(discount_group = consecutive_id(discount)) %>% summarise( discount = first(discount), start_date = first(date), end_date = last(date), price_diff = last(price) - first(price) )
运行后得到的price_diff就是你需要的每组首尾价差,不需要起止日期的话删掉对应行即可。
小提示:如果你的dplyr版本低于1.1.0,无法使用
consecutive_id(),可以把分组语句替换为group_by(discount_group = cumsum(c(TRUE, diff(discount) != 0))),效果完全一致。
方案2:base R 实现(无需额外安装包)
不想加载第三方包的话,可以用R自带的rle()函数生成连续分组ID,代码如下:
# 生成连续分组ID rle_discount <- rle(df$discount) df$discount_group <- rep(seq_along(rle_discount$lengths), rle_discount$lengths) # 按组计算价差 result <- aggregate(price ~ discount_group + discount, data = df, FUN = function(x) tail(x,1) - head(x,1))
示例数据计算结果参考
对应你给出的20行示例数据,最终会得到4个折扣组的结果:
| discount_group | discount | price_diff |
|---|---|---|
| 1 | 1 | 0.258 |
| 2 | 0 | -0.238 |
| 3 | 1 | -0.39 |
| 4 | 0 | -0.151 |
如果需要首价减末价,把计算逻辑里的last(price) - first(price)反过来即可。
内容的提问来源于stack exchange,提问作者vegatroz
相关产品推荐
相关产品推荐

