You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Pandas构建俱乐部唯一的最高总分足球阵容(含同分规则)

解决方案思路

1. 数据预处理:按优先级排序

先对每个位置的球员按总分降序、出场时间降序排序,确保每个位置里优先级最高的球员排在最前面,后续筛选时能优先锁定优质候选人。

# 假设你的DataFrame名为df,包含字段:club, position, total_score, minutes_played
df_sorted = df.sort_values(by=['position', 'total_score', 'minutes_played'], ascending=[True, False, False])

2. 定义阵型位置配额

先明确不同阵型的位置人数要求,比如:

  • 4-3-3:{'GK':1, 'DF':4, 'MF':3, 'FW':3}
  • 4-4-2:{'GK':1, 'DF':4, 'MF':4, 'FW':2}
    后续逻辑会基于这个配额筛选球员。

3. 构建最优阵容:两种可行方案

方案一:整数规划(推荐,高效准确)

用pulp库构建线性规划模型,直接求解满足约束的最优解,同时处理同分优先级:

  • 目标函数:最大化选中球员的总分之和,同时给出场时间加微小权重(比如除以10000),确保同分情况下出场时间更长的球员被优先选中
  • 约束条件:位置人数符合阵型要求、每个俱乐部最多选1人

代码示例:

import pulp

# 以4-3-3阵型为例定义配额
formation = {'GK':1, 'DF':4, 'MF':3, 'FW':3}

# 创建最大化问题
prob = pulp.LpProblem("OptimalFootballTeam", pulp.LpMaximize)

# 为每个球员创建0/1变量(1表示选中)
player_vars = pulp.LpVariable.dicts("Player", df.index, cat='Binary')

# 目标函数:总分 + 微小权重的出场时间(解决同分优先级)
prob += pulp.lpSum(
    [(df.loc[i, 'total_score'] + df.loc[i, 'minutes_played']/10000) * player_vars[i] 
     for i in df.index]
)

# 约束1:每个位置的选中人数匹配阵型
for pos, count in formation.items():
    prob += pulp.lpSum([player_vars[i] for i in df.index if df.loc[i, 'position'] == pos]) == count

# 约束2:每个俱乐部最多选中1人
for club in df['club'].unique():
    prob += pulp.lpSum([player_vars[i] for i in df.index if df.loc[i, 'club'] == club]) <= 1

# 求解(关闭日志输出)
prob.solve(pulp.PULP_CBC_CMD(msg=0))

# 获取最终选中的球员
selected_players = df.loc[[i for i in df.index if pulp.value(player_vars[i]) == 1]]
print(selected_players)

方案二:回溯法(适合小数据集)

通过递归遍历候选球员,跳过已选俱乐部的候选人,记录总分最高的组合,同时处理同分情况:

best_team = None
best_total = -float('inf')
# 定义4-3-3阵型配额
formation = {'GK':1, 'DF':4, 'MF':3, 'FW':3}

def backtrack(positions_left, selected_clubs, current_team, current_total):
    global best_team, best_total
    # 所有位置选完,更新最优解
    if not positions_left:
        if current_total > best_total:
            best_total = current_total
            best_team = current_team.copy()
        elif current_total == best_total:
            # 同分比较总出场时间
            current_minutes = sum(df.loc[p, 'minutes_played'] for p in current_team)
            best_minutes = sum(df.loc[p, 'minutes_played'] for p in best_team)
            if current_minutes > best_minutes:
                best_team = current_team.copy()
        return
    
    current_pos = next(iter(positions_left.keys()))
    needed = positions_left[current_pos]
    
    # 遍历当前位置未被选俱乐部的球员(按优先级排序)
    for idx, row in df_sorted[df_sorted['position'] == current_pos].iterrows():
        if row['club'] in selected_clubs:
            continue
        # 选择该球员
        selected_clubs.add(row['club'])
        current_team.append(idx)
        # 更新剩余位置需求
        positions_left[current_pos] -= 1
        if positions_left[current_pos] == 0:
            del positions_left[current_pos]
        # 递归
        backtrack(positions_left.copy(), selected_clubs.copy(), current_team.copy(), current_total + row['total_score'])
        # 回溯
        if current_pos not in positions_left:
            positions_left[current_pos] = 1
        else:
            positions_left[current_pos] += 1
        current_team.pop()
        selected_clubs.remove(row['club'])

# 启动回溯
backtrack(formation.copy(), set(), [], 0)
# 输出最优队伍
if best_team:
    print(df.loc[best_team])

4. 消除重复队伍的方法

重复队伍通常源于同分同优先级的球员组合,或算法遍历了相同选择路径,可通过以下方式解决:

  • 预处理阶段:对每个位置+俱乐部的球员只保留优先级最高的1位(即同位置同俱乐部仅留总分最高、出场时间最长的球员),减少候选数量
  • 回溯法中:记录已访问过的俱乐部组合,跳过重复的选择路径
  • 整数规划中:通过给出场时间加微小权重,确保最优解唯一;若仍有多个最优解,可增加额外优先级规则(如球员年龄、进球数等)

内容的提问来源于stack exchange,提问作者jscholten

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 09:24:58