Python Pandas:利用groupby标记各周期内的POI(最值点)
Correct Implementation for POI Marking per Cycle Group
Your original code has a critical mistake: you're calling Data3.groupby('Cycle') where Data3 is the raw dictionary, not the pandas DataFrame df. Groupby operations must be performed on the DataFrame object, not the source dictionary.
Corrected Apply + np.select Approach
This fixes your original approach by targeting the DataFrame for grouping, then aligning the result to the original index:
import pandas as pd import numpy as np Data3 = {'Cycle': ['1', '1', '1', '1', '1', '1', '2','2', '2', '2', '2', '2'], 'Value': [20, 24, 25, 18,15,12,1,2,19,18,12,1], 'Diff2':[1,10,18,-12,-50,-14,14,150,130,-140,12,14], } df = pd.DataFrame(Data3) df['POI'] = df.groupby('Cycle').apply( lambda g: np.select([g['Value'] == g['Value'].max(), g['Diff2'] == g['Diff2'].min()], ['A', 'B'], default='-') ).explode().reset_index(drop=True) print(df)
More Efficient Transform Approach
Using transform is cleaner and faster, as it directly returns values aligned with the original DataFrame index, eliminating the need for explode and index adjustment:
import pandas as pd import numpy as np Data3 = {'Cycle': ['1', '1', '1', '1', '1', '1', '2','2', '2', '2', '2', '2'], 'Value': [20, 24, 25, 18,15,12,1,2,19,18,12,1], 'Diff2':[1,10,18,-12,-50,-14,14,150,130,-140,12,14], } df = pd.DataFrame(Data3) # Calculate group-wise max Value and min Diff2, aligned to original rows max_value_per_cycle = df.groupby('Cycle')['Value'].transform('max') min_diff2_per_cycle = df.groupby('Cycle')['Diff2'].transform('min') # Apply conditions to set POI df['POI'] = np.select( [df['Value'] == max_value_per_cycle, df['Diff2'] == min_diff2_per_cycle], ['A', 'B'], default='-' ) print(df)
Output
Both approaches produce the target DataFrame:
Cycle Value Diff2 POI 0 1 20 1 - 1 1 24 10 - 2 1 25 18 A 3 1 18 -12 - 4 1 15 -50 B 5 1 12 -14 - 6 2 1 14 - 7 2 2 150 - 8 2 19 130 A 9 2 18 -140 B 10 2 12 12 - 11 2 1 14 -
内容的提问来源于stack exchange,提问作者SeanK22
相关产品推荐
相关产品推荐

