Pandas双列累计求和任意阈值触发同步重置及分组编号问题
实现思路
你之前的错误核心是分开计算两个列的累计和,二者重置逻辑相互独立没有绑定,只要把两个累计和的计算、阈值判断、重置逻辑放在同一个遍历流程中处理即可,任意一列触发阈值就同步重置两个累计值,同时store编号自增。
正确实现代码
import pandas as pd # 构造测试数据集 data = { 'item': ['A', 'B', 'C', 'D', 'E', 'F', 'G', 'H'], 'price': [10, 20, 5, 6, 7, 8, 20, 6], 'units': [1, 5, 8, 3, 4, 2, 2, 3] } df = pd.DataFrame(data) threshold_price = 31 threshold_units = 8 # 初始化状态变量 cur_cum_price = 0 cur_cum_units = 0 cur_store = 0 cs_price_res = [] cs_units_res = [] store_res = [] for _, row in df.iterrows(): next_price = cur_cum_price + row['price'] next_units = cur_cum_units + row['units'] # 任意阈值超限就触发同步重置 if next_price > threshold_price or next_units > threshold_units: cur_cum_price = row['price'] cur_cum_units = row['units'] cur_store += 1 else: cur_cum_price = next_price cur_cum_units = next_units # 存入结果 cs_price_res.append(cur_cum_price) cs_units_res.append(cur_cum_units) store_res.append(cur_store) # 写入DataFrame df['cs_price'] = cs_price_res df['cs_units'] = cs_units_res df['store'] = store_res
输出结果
| item | price | units | cs_price | cs_units | store |
|---|---|---|---|---|---|
| A | 10 | 1 | 10 | 1 | 0 |
| B | 20 | 5 | 30 | 6 | 0 |
| C | 5 | 8 | 5 | 8 | 1 |
| D | 6 | 3 | 6 | 3 | 2 |
| E | 7 | 4 | 13 | 7 | 2 |
| F | 8 | 2 | 8 | 2 | 3 |
| G | 20 | 2 | 28 | 4 | 3 |
| H | 6 | 3 | 6 | 3 | 4 |
内容的提问来源于stack exchange,提问作者Phoenix Islands
相关产品推荐
相关产品推荐

