如何基于Pandas构建俱乐部唯一的最高总分足球阵容(含同分规则)
解决方案思路
1. 数据预处理:按优先级排序
先对每个位置的球员按总分降序、出场时间降序排序,确保每个位置里优先级最高的球员排在最前面,后续筛选时能优先锁定优质候选人。
# 假设你的DataFrame名为df,包含字段:club, position, total_score, minutes_played df_sorted = df.sort_values(by=['position', 'total_score', 'minutes_played'], ascending=[True, False, False])
2. 定义阵型位置配额
先明确不同阵型的位置人数要求,比如:
- 4-3-3:
{'GK':1, 'DF':4, 'MF':3, 'FW':3} - 4-4-2:
{'GK':1, 'DF':4, 'MF':4, 'FW':2}
后续逻辑会基于这个配额筛选球员。
3. 构建最优阵容:两种可行方案
方案一:整数规划(推荐,高效准确)
用pulp库构建线性规划模型,直接求解满足约束的最优解,同时处理同分优先级:
- 目标函数:最大化选中球员的总分之和,同时给出场时间加微小权重(比如除以10000),确保同分情况下出场时间更长的球员被优先选中
- 约束条件:位置人数符合阵型要求、每个俱乐部最多选1人
代码示例:
import pulp # 以4-3-3阵型为例定义配额 formation = {'GK':1, 'DF':4, 'MF':3, 'FW':3} # 创建最大化问题 prob = pulp.LpProblem("OptimalFootballTeam", pulp.LpMaximize) # 为每个球员创建0/1变量(1表示选中) player_vars = pulp.LpVariable.dicts("Player", df.index, cat='Binary') # 目标函数:总分 + 微小权重的出场时间(解决同分优先级) prob += pulp.lpSum( [(df.loc[i, 'total_score'] + df.loc[i, 'minutes_played']/10000) * player_vars[i] for i in df.index] ) # 约束1:每个位置的选中人数匹配阵型 for pos, count in formation.items(): prob += pulp.lpSum([player_vars[i] for i in df.index if df.loc[i, 'position'] == pos]) == count # 约束2:每个俱乐部最多选中1人 for club in df['club'].unique(): prob += pulp.lpSum([player_vars[i] for i in df.index if df.loc[i, 'club'] == club]) <= 1 # 求解(关闭日志输出) prob.solve(pulp.PULP_CBC_CMD(msg=0)) # 获取最终选中的球员 selected_players = df.loc[[i for i in df.index if pulp.value(player_vars[i]) == 1]] print(selected_players)
方案二:回溯法(适合小数据集)
通过递归遍历候选球员,跳过已选俱乐部的候选人,记录总分最高的组合,同时处理同分情况:
best_team = None best_total = -float('inf') # 定义4-3-3阵型配额 formation = {'GK':1, 'DF':4, 'MF':3, 'FW':3} def backtrack(positions_left, selected_clubs, current_team, current_total): global best_team, best_total # 所有位置选完,更新最优解 if not positions_left: if current_total > best_total: best_total = current_total best_team = current_team.copy() elif current_total == best_total: # 同分比较总出场时间 current_minutes = sum(df.loc[p, 'minutes_played'] for p in current_team) best_minutes = sum(df.loc[p, 'minutes_played'] for p in best_team) if current_minutes > best_minutes: best_team = current_team.copy() return current_pos = next(iter(positions_left.keys())) needed = positions_left[current_pos] # 遍历当前位置未被选俱乐部的球员(按优先级排序) for idx, row in df_sorted[df_sorted['position'] == current_pos].iterrows(): if row['club'] in selected_clubs: continue # 选择该球员 selected_clubs.add(row['club']) current_team.append(idx) # 更新剩余位置需求 positions_left[current_pos] -= 1 if positions_left[current_pos] == 0: del positions_left[current_pos] # 递归 backtrack(positions_left.copy(), selected_clubs.copy(), current_team.copy(), current_total + row['total_score']) # 回溯 if current_pos not in positions_left: positions_left[current_pos] = 1 else: positions_left[current_pos] += 1 current_team.pop() selected_clubs.remove(row['club']) # 启动回溯 backtrack(formation.copy(), set(), [], 0) # 输出最优队伍 if best_team: print(df.loc[best_team])
4. 消除重复队伍的方法
重复队伍通常源于同分同优先级的球员组合,或算法遍历了相同选择路径,可通过以下方式解决:
- 预处理阶段:对每个位置+俱乐部的球员只保留优先级最高的1位(即同位置同俱乐部仅留总分最高、出场时间最长的球员),减少候选数量
- 回溯法中:记录已访问过的俱乐部组合,跳过重复的选择路径
- 整数规划中:通过给出场时间加微小权重,确保最优解唯一;若仍有多个最优解,可增加额外优先级规则(如球员年龄、进球数等)
内容的提问来源于stack exchange,提问作者jscholten
相关产品推荐
相关产品推荐

