基于PuLP的Python梦幻阵容优化器:耗时与重复阵容问题
梦幻阵容优化器:性能优化与去重解决方案
一、优化耗时问题解决方案
以下是针对代码运行效率的具体优化措施:
1. 切换为迭代式阵容生成(核心优化)
当前代码一次性创建所有阵容的变量和约束,导致变量数=球员数×阵容数,阵容数增加时变量会急剧膨胀。改为每次生成1套最优阵容,添加约束排除该阵容后,再生成下一套,变量数固定为球员数,大幅降低求解压力:
# 替换原有的多阵容变量定义和循环约束代码 prob = LpProblem("NBA_optimizer", sense=LpMaximize) # 单阵容变量 choices = LpVariable.dicts("choice", range(len(df)), cat=LpBinary) # 定义基础约束(只定义一次) # 8人阵容约束 prob += lpSum([choices[i] for i in range(len(df))]) == 8 # 薪资约束 prob += lpSum([choices[i] * df.Credits.iloc[i] for i in range(len(df))]) <= 100 # 位置约束(预定义位置过滤) pos_constraints = { "PG": (1, 4), "SG": (1, 4), "SF": (1, 4), "PF": (1, 4), "C": (1, 4) } for pos, (min_count, max_count) in pos_constraints.items(): pos_indices = df[df.Pos == pos].index.tolist() prob += lpSum([choices[i] for i in pos_indices]) >= min_count prob += lpSum([choices[i] for i in pos_indices]) <= max_count # 单球队人数约束(原代码仅限制teams[0],如需全球队约束保留循环,否则只保留目标球队) for team in teams: team_indices = df[df.Team == team].index.tolist() prob += lpSum([choices[i] for i in team_indices]) >= 3 prob += lpSum([choices[i] for i in team_indices]) <= 5 # 迭代生成阵容 lineup_list = [] start = datetime.now() for _ in range(num_lineups): # 求解当前最优阵容,启用暖启动加速 prob.solve(PULP_CBC_CMD(msg=0, timeLimit=600, threads=10, warmStart=True)) if prob.status != LpStatusOptimal: print(f"仅生成{len(lineup_list)}套阵容后无法找到最优解") break # 提取当前阵容 selected = [i for i in range(len(df)) if choices[i].varValue == 1] lineup_players = df.iloc[selected].apply(lambda x: f"({x.Pos})-{x.Name}", axis=1).tolist() lineup_list.append(sort_players(lineup_players)) # 添加约束排除当前阵容(至少1个球员不同) prob += lpSum([choices[i] for i in selected]) <= 7 # 处理Exposure约束:统计已生成阵容中球员入选次数,若未达要求可动态调整约束 player_counts = pd.Series([sum(1 for lineup in lineup_list if player in lineup) for player in df.apply(lambda x: f"({x.Pos})-{x.Name}", axis=1)]) for idx, required in props.items(): if player_counts.iloc[idx] < required: # 剩余需要生成的阵容数中补全Exposure要求 remaining = num_lineups - len(lineup_list) if remaining > 0: prob += lpSum([choices[idx] for _ in range(remaining)]) >= (required - player_counts.iloc[idx])
2. 简化约束定义,减少冗余计算
- 删除不必要的
proj和sal字典,直接调用DataFrame列数据,避免重复存储 - 预计算位置、球队的索引列表,避免每次循环重新生成
list(df.Pos == "PG") - 直接用
lpSum构建约束,无需额外创建LpAffineExpression,减少中间对象开销
3. 求解器参数调优
- 启用
warmStart=True:迭代求解时复用上一次的解作为初始值,大幅缩短求解时间 - 设置
gapRel参数:若无需绝对最优解,允许一定的最优间隙(如gapRel=0.01),提前终止求解 - 将
threads设为CPU核心数(如threads=os.cpu_count()),最大化利用硬件资源
二、重复阵容问题解决方案
当前代码通过比较阵容投影和的方式防重复,存在漏洞:不同球员组合可能有相同的投影和,导致重复阵容。以下是可靠的解决方法:
1. 添加排除已有阵容的硬约束
在每次生成一套阵容后,添加约束:新阵容最多只能包含已有阵容中的7名球员(确保至少1名球员不同),代码示例见上述迭代式生成部分的prob += lpSum([choices[i] for i in selected]) <= 7。
2. 事后验证并去重(可选)
生成所有阵容后,可通过集合去重验证唯一性:
# 验证并去重阵容 unique_lineups = [] seen = set() for lineup in lineup_list: lineup_tuple = tuple(sorted(lineup)) if lineup_tuple not in seen: seen.add(lineup_tuple) unique_lineups.append(lineup) if len(unique_lineups) != num_lineups: print(f"目标生成{num_lineups}套阵容,实际唯一阵容{len(unique_lineups)}套") lineup_list = unique_lineups
3. 修复一次性生成模式的去重逻辑(若坚持原架构)
如果必须一次性生成所有阵容,替换原有的投影和比较约束,改为针对每对阵容添加差异约束:
# 替换原j>0时的去重约束 for j in range(num_lineups): # ... 其他约束 ... for k in range(j): # 约束阵容j和k至少有一个球员不同 prob += lpSum([choices[i,j] - choices[i,k] for i in range(len(df))]) >= 1 prob += lpSum([choices[i,k] - choices[i,j] for i in range(len(df))]) >= 1
注意:这种方式会导致约束数呈平方级增长,仅适合小数量阵容场景。
内容的提问来源于stack exchange,提问作者Yagami_Light
相关产品推荐
相关产品推荐

