如何用Python函数实现DataFrame按最高销量筛选并删除指定行?
问题
places_within_catchment字段存储着place_id的列表,处理逻辑如下:
- 按原始行顺序遍历,对每一行,从其
places_within_catchment列表中找到当前仍存在的place_id里销量最高的那个。 - 删除该列表中除这个最高销量place_id之外的所有id对应的行。
- 已被删除的行跳过处理。
举个例子:
- 处理第一行时,其
places_within_catchment是[2,3],其中3的销量最高,因此删除所有place_id=2的行。 - 接着处理原第三行(原第二行已被删除),其
places_within_catchment是[1,2,5],此时2已不存在,剩余1和5中1的销量更高,因此删除place_id=5的行。
请问如何用Python函数实现该逻辑?
我尝试的代码
result_df = pd.DataFrame(columns=new_df_copy.columns) for index, row in new_df_copy.iterrows(): max_avg_sales = 0 max_id = 0 # Find the id with the highest avg_sales within places_within_catchment for place_id in row['places_within_catchment']: if place_id in new_df_copy['id'].values: avg_sales = new_df_copy.loc[new_df_copy['id'] == place_id, 'avg_sales'].values[0] if avg_sales > max_avg_sales: max_avg_sales = avg_sales max_id = place_id # Append the row with the maximum avg_sales to the result DataFrame result_df = pd.concat([result_df, new_df_copy.loc[new_df_copy['id'] == max_id]], ignore_index=True) # Display the result DataFrame print(result_df)
示例数据生成代码
data = { 'place_id': list(range(1, 6)), 'avg_sales': [500.4, 200.4, 600.25, 200.93, 60.1], 'places_within_catchment': [[2, 3], [1, 3, 4, 5], [1, 2, 5], [1], [1, 3]] } new_df = pd.DataFrame(data) print(new_df)
示例表格:
预期输出逻辑说明:
解决方案
下面的函数严格按照需求实现了逐步删除逻辑:
import pandas as pd def process_catchment(df): # 复制原始数据作为工作副本,避免修改原数据 working_df = df.copy() # 按原始place_id顺序遍历 original_ids = df['place_id'].tolist() for pid in original_ids: # 如果当前id已被删除,跳过 if pid not in working_df['place_id'].values: continue # 获取当前行的catchment列表 current_catchment = working_df.loc[working_df['place_id'] == pid, 'places_within_catchment'].iloc[0] # 筛选出当前仍在工作副本中的catchment id existing_catchment = [cid for cid in current_catchment if cid in working_df['place_id'].values] if not existing_catchment: continue # 找到当前catchment中销量最高的id max_row = working_df[working_df['place_id'].isin(existing_catchment)].nlargest(1, 'avg_sales') max_id = max_row['place_id'].iloc[0] # 确定要删除的id:catchment中除了max_id之外的所有存在的id to_delete = [cid for cid in existing_catchment if cid != max_id] # 删除对应行 working_df = working_df[~working_df['place_id'].isin(to_delete)] return working_df # 测试示例 if __name__ == "__main__": data = { 'place_id': list(range(1, 6)), 'avg_sales': [500.4, 200.4, 600.25, 200.93, 60.1], 'places_within_catchment': [[2, 3], [1, 3, 4, 5], [1, 2, 5], [1], [1, 3]] } df = pd.DataFrame(data) final_df = process_catchment(df) print("处理后结果:") print(final_df)
代码说明
- 数据副本:使用
df.copy()创建工作副本,防止修改原始数据。 - 顺序遍历:按原始数据的place_id顺序处理,跳过已被删除的id。
- 筛选有效id:对每行的catchment列表,只保留当前仍存在于工作副本中的id。
- 找最高销量id:用
nlargest快速找到销量最高的id。 - 删除指定行:移除catchment列表中除最高销量id外的所有id对应的行,逐步精简数据。
运行后会得到预期结果:
place_id avg_sales places_within_catchment 0 1 500.40 [2, 3] 1 3 600.25 [1, 2, 5] 2 4 200.93 [1]
内容的提问来源于stack exchange,提问作者karan
相关产品推荐
相关产品推荐

