Pandas遍历DataFrame行实现带中断条件的匹配结果提取
解决方案
核心逻辑
- 按
ID分组,仅处理同ID下的记录 - 每组内按
MT升序排序,保证遍历顺序是从MT更小的历史行到当前行 - 对每行遍历其之前的历史行:
- 若历史行
Price小于当前行Price,收集对应的Price和Date - 一旦遇到
Price大于等于当前行的记录,立即中断遍历
- 若历史行
代码实现
先给出示例数据(可替换成你实际的DataFrame):
import pandas as pd data = { 'ID': ['A', 'A', 'A', 'B', 'B', 'B'], 'MT': [1, 2, 3, 1, 2, 3], 'Price': [10, 8, 12, 15, 13, 16], 'Date': ['2023-01-01', '2023-01-02', '2023-01-03', '2023-01-01', '2023-01-02', '2023-01-03'] } df = pd.DataFrame(data)
完善后的处理代码:
def process_single_id(group): # 按MT升序排序,确保历史行在前、当前行在后 sorted_group = group.sort_values('MT').reset_index(drop=True) matched_prices = [] matched_dates = [] for idx in range(len(sorted_group)): current_price = sorted_group.loc[idx, 'Price'] temp_p = [] temp_d = [] # 遍历当前行之前的所有历史行 for hist_idx in range(idx): hist_price = sorted_group.loc[hist_idx, 'Price'] if hist_price < current_price: temp_p.append(hist_price) temp_d.append(sorted_group.loc[hist_idx, 'Date']) else: # 遇到不符合条件的记录,直接中断遍历 break matched_prices.append(temp_p) matched_dates.append(temp_d) sorted_group['Matched_Price'] = matched_prices sorted_group['Date_Values'] = matched_dates return sorted_group # 分组处理后合并结果 final_df = df.groupby('ID').apply(process_single_id).reset_index(drop=True) print(final_df)
代码说明
process_single_id函数专门处理单个ID的分组数据:- 先对分组内数据按
MT排序,保证历史记录的顺序正确 - 外层循环遍历每行,内层循环检查该行之前的所有历史行
- 内层循环中用
break实现中断逻辑,一旦遇到Price不小于当前值的记录,立即停止后续检查
- 先对分组内数据按
- 最后通过
groupby+apply完成所有ID分组的处理,合并得到最终结果
示例输出
运行代码后会得到如下结果:
ID MT Price Date Matched_Price Date_Values 0 A 1 10 2023-01-01 [] [] 1 A 2 8 2023-01-02 [] [] 2 A 3 12 2023-01-03 [10, 8] [2023-01-01, 2023-01-02] 3 B 1 15 2023-01-01 [] [] 4 B 2 13 2023-01-02 [] [] 5 B 3 16 2023-01-03 [15, 13] [2023-01-01, 2023-01-02]
内容的提问来源于stack exchange,提问作者Gopinathan
相关产品推荐
相关产品推荐

