You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pokerenv库报错TypeError: unsupported operand type(s) for >>: 'list' and 'int'

问题:使用pokerenv运行强化学习示例时出现随机类型错误

在强化学习项目中使用pokerenv库运行官方示例代码时,会随机触发类型错误:unsupported operand type(s) for >>: 'list' and 'int'。错误并非立即出现,通常在环境运行10-20步后随机触发。

报错栈显示问题根源在treys库的Card.get_rank_int方法——该方法预期接收int类型参数,但实际传入了list类型。查看pokerenv源码后发现,self.cards被初始化为列表,而self.deck.draw(1)的返回值是列表(而非单个整数),直接追加到self.cards后,后续调用Card.int_to_str时传入列表元素,导致类型错误。

完整报错信息

---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
Cell In[7], line 13
     11 while True:
     12     action = agents[acting_player].get_action(obs)
---> 13     print(f"acion{table.step(action)} tipo{table.step(action)}")
     14     obs, reward, done, _ = table.step(action)
     15     print(f"obs={obs} tipo{type(obs)}")

File /opt/conda/lib/python3.10/site-packages/pokerenv/table.py:199, in Table.step(self, action)
    196                 self.next_player_i = min(active_players_before)
    198 if self.street_finished and not self.hand_is_over:
---> 199     self._street_transition()
    201 obs = np.zeros(self.observation_space.shape[0]) if self.hand_is_over else self._get_observation(self.players[self.next_player_i])
    202 rewards = np.asarray([player.get_reward() for player in sorted(self.players)])

File /opt/conda/lib/python3.10/site-packages/pokerenv/table.py:219, in Table._street_transition(self, transition_to_end)
    215 new = self.deck.draw(1)
    216 self.cards.append(new)
    217 self._write_event("*** TURN *** [%s %s %s] [%s]" %
    218                   (Card.int_to_str(self.cards[0]), Card.int_to_str(self.cards[1]),
---> 219                    Card.int_to_str(self.cards[2]), Card.int_to_str(self.cards[3])))
    220 self.street = GameState.TURN
    221 transitioned = True

File /opt/conda/lib/python3.10/site-packages/treys/card.py:81, in Card.int_to_str(card_int)
     79 @staticmethod
     80 def int_to_str(card_int: int) -> str:
---> 81     rank_int = Card.get_rank_int(card_int)
     82     suit_int = Card.get_suit_int(card_int)
     83     return Card.STR_RANKS[rank_int] + Card.INT_SUIT_TO_CHAR_SUIT[suit_int]

File /opt/conda/lib/python3.10/site-packages/treys/card.py:87, in Card.get_rank_int(card_int)
     85 @staticmethod
     86 def get_rank_int(card_int: int) -> int:
---> 87     return (card_int >> 8) & 0xF

TypeError: unsupported operand type(s) for >>: 'list' and 'int'

示例代码

import numpy as np
import pokerenv.obs_indices as indices
from pokerenv.table import Table
from pokerenv.common import PlayerAction, Action, action_list


class ExampleRandomAgent:
    def __init__(self):
        self.actions = []
        self.observations = []
        self.rewards = []

    def get_action(self, observation):
        self.observations.append(observation)
        valid_actions = np.argwhere(observation[indices.VALID_ACTIONS] == 1).flatten()
        valid_bet_low = observation[indices.VALID_BET_LOW]
        valid_bet_high = observation[indices.VALID_BET_HIGH]
        chosen_action = PlayerAction(np.random.choice(valid_actions))
        bet_size = 0
        if chosen_action is PlayerAction.BET:
            bet_size = np.random.uniform(valid_bet_low, valid_bet_high)
        table_action = Action(chosen_action, bet_size)
        self.actions.append(table_action)
        return table_action

    def reset(self):
        self.actions = []
        self.observations = []
        self.rewards = []

active_players = 6
agents = [ExampleRandomAgent() for _ in range(6)]
player_names = {0: 'TrackedAgent1', 1: 'Agent2'} # Rest are defaulted to player3, player4...
# Should we only log the 0th players (here TrackedAgent1) private cards to hand history files
track_single_player = True 
# Bounds for randomizing player stack sizes in reset()
low_stack_bbs = 50
high_stack_bbs = 200
hand_history_location = 'hands/'
invalid_action_penalty = 0
table = Table(active_players, 
              player_names=player_names,
              track_single_player=track_single_player,
              stack_low=low_stack_bbs,
              stack_high=high_stack_bbs,
              hand_history_location=hand_history_location,
              invalid_action_penalty=invalid_action_penalty
)
table.seed(1)

iteration = 1
while True:
    if iteration % 50 == 0:
        table.hand_history_enabled = True
    active_players = np.random.randint(2, 7)
    table.n_players = active_players
    obs = table.reset()
    for agent in agents:
        agent.reset()
    acting_player = int(obs[indices.ACTING_PLAYER])
    while True:
        action = agents[acting_player].get_action(obs)
        obs, reward, done, _ = table.step(action)
        #print(f"obs={obs} tipo{type(obs)}")
        #print(f"reward={reward} tipo{type(reward)}")
        #print(f"done={done} tipo{type(done)}")
        if  done:
            # Distribute final rewards
            for i in range(active_players):
                agents[i].rewards.append(reward[i])
            break
        else:
            # This step can be skipped unless invalid action penalty is enabled, 
            # since we only get a reward when the pot is distributed, and the done flag is set
            agents[acting_player].rewards.append(reward[acting_player])
            acting_player = int(obs[indices.ACTING_PLAYER])
    iteration += 1
    table.hand_history_enabled = False

解决方案

1. 修复pokerenv源码中的卡牌存储逻辑

找到pokerenv/table.py中的_street_transition方法,修改卡牌追加代码:
原代码:

new = self.deck.draw(1)
self.cards.append(new)

修改为:

new = self.deck.draw(1)[0]  # 取出列表中的单个卡牌整数
self.cards.append(new)

这样self.cards中的元素都是int类型,后续调用Card.int_to_str时就不会触发类型错误。

2. 避免重复调用table.step(action)

注意到示例代码中曾有重复调用table.step(action)的代码(已注释),这会导致环境状态被推进两次,可能引发异常行为。确保每个action只调用一次step方法,保持环境状态的一致性。

内容的提问来源于stack exchange,提问作者darth momin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 18:45:55