基于规则的推荐:用Python实现带菜系限制的好友餐厅Top4推荐
问题描述
现有一个包含3位好友、10家餐厅的DataFrame,数据包含好友姓名、餐厅名称、兴趣排名(InterestRank,1表示最感兴趣,10表示最不感兴趣),以及餐厅的菜系(Cuisine)、消费档次(Cost)、是否提供酒精饮品(Alcohol)等属性,原始数据如下:
FriendName,Restaurant,InterestRank,Cuisine,Cost,Alcohol Amy,R2,1,French,$$,No Ben,R2,3,French,$$,No Cathy,R2,8,French,$$,No Amy,R1,2,French,$$$,Yes Ben,R1,9,French,$$$,Yes Cathy,R1,5,French,$$$,Yes Amy,R4,3,French,$$$,Yes Ben,R4,5,French,$$$,Yes Cathy,R4,10,French,$$$,Yes Amy,R3,4,French,$$,Yes Ben,R3,10,French,$$,Yes Cathy,R3,6,French,$$,Yes Amy,R10,5,Mexican,$$$,Yes Ben,R10,6,Mexican,$$$,Yes Cathy,R10,7,Mexican,$$$,Yes Amy,R7,6,Japanese,$$,Yes Ben,R7,1,Japanese,$$,Yes Cathy,R7,9,Japanese,$$,Yes Amy,R6,7,Japanese,$,No Ben,R6,8,Japanese,$,No Cathy,R6,3,Japanese,$,No Amy,R8,8,Mexican,$$,No Ben,R8,4,Mexican,$$,No Cathy,R8,2,Mexican,$$,No Amy,R5,9,Japanese,$$,No Ben,R5,2,Japanese,$$,No Cathy,R5,1,Japanese,$$,No Amy,R9,10,Mexican,$$,No Ben,R9,7,Mexican,$$,No Cathy,R9,4,Mexican,$$,No
需要为每位好友推荐Top4餐厅,要求:
- 按InterestRank升序排序(数值越小越优先)
- 同一菜系的推荐数量不超过2家
预期输出的DataFrame如下:
| FriendName | Restaurant | RecommendationRank |
|---|---|---|
| Amy | R2 | 1 |
| Amy | R1 | 2 |
| Amy | R10 | 3 |
| Amy | R7 | 4 |
| Ben | R7 | 1 |
| Ben | R2 | 2 |
| Ben | R5 | 3 |
| Ben | R8 | 4 |
| Cathy | R5 | 1 |
| Cathy | R8 | 2 |
| Cathy | R6 | 3 |
| Cathy | R9 | 4 |
Pythonic实现方案
利用Pandas的分组功能结合自定义筛选逻辑即可实现需求,核心思路是按好友分组后,对每个好友的餐厅列表先按兴趣排名排序,再逐个筛选并控制每个菜系的入选数量,直到选够4家为止。
import pandas as pd # 构造原始DataFrame(也可通过pd.read_csv读取本地CSV文件) data = [ ["Amy", "R2", 1, "French", "$$", "No"], ["Ben", "R2", 3, "French", "$$", "No"], ["Cathy", "R2", 8, "French", "$$", "No"], ["Amy", "R1", 2, "French", "$$$", "Yes"], ["Ben", "R1", 9, "French", "$$$", "Yes"], ["Cathy", "R1", 5, "French", "$$$", "Yes"], ["Amy", "R4", 3, "French", "$$$", "Yes"], ["Ben", "R4", 5, "French", "$$$", "Yes"], ["Cathy", "R4", 10, "French", "$$$", "Yes"], ["Amy", "R3", 4, "French", "$$", "Yes"], ["Ben", "R3", 10, "French", "$$", "Yes"], ["Cathy", "R3", 6, "French", "$$", "Yes"], ["Amy", "R10", 5, "Mexican", "$$$", "Yes"], ["Ben", "R10", 6, "Mexican", "$$$", "Yes"], ["Cathy", "R10", 7, "Mexican", "$$$", "Yes"], ["Amy", "R7", 6, "Japanese", "$$", "Yes"], ["Ben", "R7", 1, "Japanese", "$$", "Yes"], ["Cathy", "R7", 9, "Japanese", "$$", "Yes"], ["Amy", "R6", 7, "Japanese", "$", "No"], ["Ben", "R6", 8, "Japanese", "$", "No"], ["Cathy", "R6", 3, "Japanese", "$", "No"], ["Amy", "R8", 8, "Mexican", "$$", "No"], ["Ben", "R8", 4, "Mexican", "$$", "No"], ["Cathy", "R8", 2, "Mexican", "$$", "No"], ["Amy", "R5", 9, "Japanese", "$$", "No"], ["Ben", "R5", 2, "Japanese", "$$", "No"], ["Cathy", "R5", 1, "Japanese", "$$", "No"], ["Amy", "R9", 10, "Mexican", "$$", "No"], ["Ben", "R9", 7, "Mexican", "$$", "No"], ["Cathy", "R9", 4, "Mexican", "$$", "No"], ] df = pd.DataFrame(data, columns=["FriendName", "Restaurant", "InterestRank", "Cuisine", "Cost", "Alcohol"]) def get_top_recommendations(group): # 按兴趣排名升序排序,优先推荐兴趣度高的餐厅 sorted_group = group.sort_values("InterestRank", ascending=True) recommendations = [] cuisine_count = {} for _, row in sorted_group.iterrows(): cuisine = row["Cuisine"] # 若当前菜系已选2家,跳过该餐厅 if cuisine_count.get(cuisine, 0) >= 2: continue recommendations.append(row) cuisine_count[cuisine] = cuisine_count.get(cuisine, 0) + 1 # 选够4家后停止筛选 if len(recommendations) == 4: break # 转换为DataFrame并添加推荐排名 rec_df = pd.DataFrame(recommendations) rec_df["RecommendationRank"] = range(1, len(rec_df)+1) return rec_df[["FriendName", "Restaurant", "RecommendationRank"]] # 按好友分组处理,合并结果后重置索引 result = df.groupby("FriendName").apply(get_top_recommendations).reset_index(drop=True) print(result)
代码说明
- 数据构造/读取:可根据实际场景选择直接构造DataFrame或读取本地CSV文件。
- 分组逻辑:通过
groupby("FriendName")将数据按好友拆分,对每个分组应用自定义筛选函数。 - 排序与筛选:每个好友的餐厅列表先按InterestRank升序排序,遍历过程中统计每个菜系的入选数量,确保单菜系不超过2家,直到选满4家。
- 生成推荐排名:为选中的餐厅添加1到4的推荐排名,最后保留需求指定的列输出。
内容的提问来源于stack exchange,提问作者Utsav Maniar
相关产品推荐
相关产品推荐

