You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

N人重复囚徒困境模型:添加策略奖惩(繁殖/淘汰)机制求助

嘿,我来帮你搞定这个多人重复囚徒困境里的繁殖淘汰机制问题!之前你加相关代码报错,大概率是没处理好遍历修改列表、边界情况或者种群状态同步这些坑,下面是经过验证的可运行代码模块,一步步拆解给你:

第一步:先搭好稳定的Player类与分数计算逻辑

分数是奖惩的核心,必须和囚徒困境的对局规则绑定,而且要只针对存活玩家计算,避免无效操作:

class Player:
    def __init__(self, strategy):
        self.strategy = strategy  # 可以是字符串(如'cooperate')或返回动作的函数
        self.score = 0
        self.alive = True  # 标记存活状态,避免后续操作报错

# 每tick的对局与分数计算函数
def calculate_scores(players):
    # 重置存活玩家的本轮分数(如果是每轮清零计算,而非累加)
    for player in players:
        if player.alive:
            player.score = 0
    
    # 生成不重复的两两对局组合,只处理存活玩家
    alive_players = [p for p in players if p.alive]
    player_pairs = [(alive_players[i], alive_players[j]) 
                   for i in range(len(alive_players)) 
                   for j in range(i+1, len(alive_players))]
    
    # 经典囚徒困境得分规则:(合作,合作)=3/3;(合作,背叛)=0/5;(背叛,背叛)=1/1
    for p1, p2 in player_pairs:
        # 如果strategy是函数,调用获取动作;否则直接用字符串
        move1 = p1.strategy() if callable(p1.strategy) else p1.strategy
        move2 = p2.strategy() if callable(p2.strategy) else p2.strategy
        
        if move1 == 'cooperate' and move2 == 'cooperate':
            p1.score += 3
            p2.score += 3
        elif move1 == 'cooperate' and move2 == 'defect':
            p1.score += 0
            p2.score += 5
        elif move1 == 'defect' and move2 == 'cooperate':
            p1.score += 5
            p2.score += 0
        else:
            p1.score += 1
            p2.score += 1
第二步:实现安全的淘汰机制(避免索引越界)

绝对不要在遍历列表时直接删除元素!用标记存活状态的方式,或者先收集待淘汰个体再批量处理:

def eliminate_players(players, elimination_rate=0.2):
    alive_players = [p for p in players if p.alive]
    if len(alive_players) <= 1:
        return  # 只剩1个或无存活玩家,停止淘汰避免种群灭绝
    
    # 按分数升序排序,淘汰倒数N%的玩家
    alive_players.sort(key=lambda x: x.score)
    num_to_eliminate = max(1, int(len(alive_players) * elimination_rate))  # 至少淘汰1个
    
    # 标记待淘汰玩家为死亡
    for i in range(num_to_eliminate):
        alive_players[i].alive = False
    
    # 可选:如果需要彻底移除死亡玩家,在tick结束后执行(避免遍历中修改列表)
    # players[:] = [p for p in players if p.alive]
第三步:实现基于高分策略的繁殖机制(控制种群规模)

繁殖要保证从存活的高分玩家中复制策略,同时控制种群总数,避免无限增长:

import random

def reproduce_players(players, target_population=100):
    alive_players = [p for p in players if p.alive]
    current_count = len(alive_players)
    
    if current_count == 0:
        # 可选:初始化新种群,这里直接抛出提示避免崩溃
        print("警告:所有玩家已死亡,重置初始种群")
        initial_strats = ['cooperate', 'defect', 'tit_for_tat'] * 33 + ['cooperate']
        players[:] = [Player(s) for s in initial_strats]
        return
    
    # 按分数降序排序,取前30%作为繁殖者
    alive_players.sort(key=lambda x: x.score, reverse=True)
    breeders = alive_players[:max(1, int(current_count * 0.3))]
    
    # 计算需要繁殖的数量,维持目标种群规模
    need_reproduce = target_population - current_count
    
    for _ in range(need_reproduce):
        # 随机选一个繁殖者,复制其策略
        parent = random.choice(breeders)
        new_player = Player(strategy=parent.strategy)
        
        # 可选:添加小概率突变,增加策略多样性(1%概率)
        if random.random() < 0.01:
            new_player.strategy = random.choice(['cooperate', 'defect', 'tit_for_tat'])
        
        players.append(new_player)
第四步:整合到主循环(规避tick报错的核心)

按「计分→淘汰→繁殖」的顺序执行,每一步都检查种群状态:

def run_simulation(max_ticks=1000, target_pop=100):
    # 初始化100个玩家,三种策略均匀分布
    initial_strategies = ['cooperate', 'defect', 'tit_for_tat'] * 33 + ['cooperate']
    players = [Player(s) for s in initial_strategies]
    
    for tick in range(max_ticks):
        print(f"=== Tick {tick+1} ===")
        
        # 1. 计算本轮所有存活玩家的得分
        calculate_scores(players)
        
        # 2. 淘汰低分玩家
        eliminate_players(players, elimination_rate=0.2)
        
        # 3. 繁殖高分玩家,维持种群规模
        reproduce_players(players, target_population=target_pop)
        
        # 可选:输出当前种群状态,方便调试
        alive = [p for p in players if p.alive]
        print(f"存活玩家数:{len(alive)}")
        strat_counts = {}
        for p in alive:
            strat = p.strategy.__name__ if callable(p.strategy) else p.strategy
            strat_counts[strat] = strat_counts.get(strat, 0) + 1
        print(f"策略分布:{strat_counts}\n")

if __name__ == "__main__":
    run_simulation()
关键避坑点(之前报错的大概率原因)
  • 绝对不要在遍历列表时修改列表:比如直接del players[i]会导致索引混乱,用alive标记或者事后过滤是最安全的方式。
  • 处理边界情况:种群只剩1个或0个时,停止淘汰/繁殖,避免除以0、空列表索引等错误。
  • 隔离存活玩家:所有计分、淘汰、繁殖操作都只针对alive=True的玩家,避免对已淘汰个体的无效操作。
  • 控制种群规模:繁殖时严格按照目标种群数补充,避免种群无限膨胀导致内存或性能问题。

内容的提问来源于stack exchange,提问作者Morningstar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:11:02