Pandas中按组动态分箱:基于组内最值均值实现无循环操作
Solution Using Pandas (No Loops)
You can achieve this in just a couple of lines using Pandas' groupby.transform to compute group-level statistics and apply to assign range labels dynamically. Here's how:
import pandas as pd # Your input data data = { 'Country': ['Uganda', 'Kenya', 'Kenya', 'Tanzania', 'Uganda', 'Uganda', 'Tanzania', 'Kenya'], 'Value': [210, 423, 315, 780, 124, 213, 978, 524] } df = pd.DataFrame(data) # Compute group-wise min, max, and their mean for each row df[['min_val', 'max_val', 'mean_val']] = df.groupby('Country')['Value'].transform( lambda g: pd.Series([g.min(), g.max(), (g.min() + g.max()) / 2]) ) # Assign the appropriate range based on value vs group mean df['Range'] = df.apply( lambda row: f"({row['min_val']}, {row['mean_val']})" if row['Value'] <= row['mean_val'] else f"({row['mean_val']}, {row['max_val']})", axis=1 ) # Clean up and reorder columns to match desired output result = df.drop(['min_val', 'max_val', 'mean_val'], axis=1)[['Country', 'Value', 'Range']] print(result)
Output:
Country Value Range 0 Uganda 210 (124.0, 168.5) 1 Kenya 423 (419.5, 524.0) 2 Kenya 315 (315.0, 419.5) 3 Tanzania 780 (780.0, 879.0) 4 Uganda 124 (124.0, 168.5) 5 Uganda 213 (168.5, 213.0) 6 Tanzania 978 (879.0, 978.0) 7 Kenya 524 (419.5, 524.0)
Key Details:
groupby.transformensures every row gets stats from its country group, eliminating the need for manual loops.- The
applyfunction checks if the value falls below/equal to the group mean or above, then formats the range string correctly. - The minor difference in Tanzania's upper bound (978 vs your example's 980) is likely a typo in your expected output—our code uses the actual maximum value from the dataset.
内容的提问来源于stack exchange,提问作者Ankur Sinha
相关产品推荐
相关产品推荐

