pokerenv库报错TypeError: unsupported operand type(s) for >>: 'list' and 'int'
问题:使用pokerenv运行强化学习示例时出现随机类型错误
在强化学习项目中使用pokerenv库运行官方示例代码时,会随机触发类型错误:unsupported operand type(s) for >>: 'list' and 'int'。错误并非立即出现,通常在环境运行10-20步后随机触发。
报错栈显示问题根源在treys库的Card.get_rank_int方法——该方法预期接收int类型参数,但实际传入了list类型。查看pokerenv源码后发现,self.cards被初始化为列表,而self.deck.draw(1)的返回值是列表(而非单个整数),直接追加到self.cards后,后续调用Card.int_to_str时传入列表元素,导致类型错误。
完整报错信息
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) Cell In[7], line 13 11 while True: 12 action = agents[acting_player].get_action(obs) ---> 13 print(f"acion{table.step(action)} tipo{table.step(action)}") 14 obs, reward, done, _ = table.step(action) 15 print(f"obs={obs} tipo{type(obs)}") File /opt/conda/lib/python3.10/site-packages/pokerenv/table.py:199, in Table.step(self, action) 196 self.next_player_i = min(active_players_before) 198 if self.street_finished and not self.hand_is_over: ---> 199 self._street_transition() 201 obs = np.zeros(self.observation_space.shape[0]) if self.hand_is_over else self._get_observation(self.players[self.next_player_i]) 202 rewards = np.asarray([player.get_reward() for player in sorted(self.players)]) File /opt/conda/lib/python3.10/site-packages/pokerenv/table.py:219, in Table._street_transition(self, transition_to_end) 215 new = self.deck.draw(1) 216 self.cards.append(new) 217 self._write_event("*** TURN *** [%s %s %s] [%s]" % 218 (Card.int_to_str(self.cards[0]), Card.int_to_str(self.cards[1]), ---> 219 Card.int_to_str(self.cards[2]), Card.int_to_str(self.cards[3]))) 220 self.street = GameState.TURN 221 transitioned = True File /opt/conda/lib/python3.10/site-packages/treys/card.py:81, in Card.int_to_str(card_int) 79 @staticmethod 80 def int_to_str(card_int: int) -> str: ---> 81 rank_int = Card.get_rank_int(card_int) 82 suit_int = Card.get_suit_int(card_int) 83 return Card.STR_RANKS[rank_int] + Card.INT_SUIT_TO_CHAR_SUIT[suit_int] File /opt/conda/lib/python3.10/site-packages/treys/card.py:87, in Card.get_rank_int(card_int) 85 @staticmethod 86 def get_rank_int(card_int: int) -> int: ---> 87 return (card_int >> 8) & 0xF TypeError: unsupported operand type(s) for >>: 'list' and 'int'
示例代码
import numpy as np import pokerenv.obs_indices as indices from pokerenv.table import Table from pokerenv.common import PlayerAction, Action, action_list class ExampleRandomAgent: def __init__(self): self.actions = [] self.observations = [] self.rewards = [] def get_action(self, observation): self.observations.append(observation) valid_actions = np.argwhere(observation[indices.VALID_ACTIONS] == 1).flatten() valid_bet_low = observation[indices.VALID_BET_LOW] valid_bet_high = observation[indices.VALID_BET_HIGH] chosen_action = PlayerAction(np.random.choice(valid_actions)) bet_size = 0 if chosen_action is PlayerAction.BET: bet_size = np.random.uniform(valid_bet_low, valid_bet_high) table_action = Action(chosen_action, bet_size) self.actions.append(table_action) return table_action def reset(self): self.actions = [] self.observations = [] self.rewards = [] active_players = 6 agents = [ExampleRandomAgent() for _ in range(6)] player_names = {0: 'TrackedAgent1', 1: 'Agent2'} # Rest are defaulted to player3, player4... # Should we only log the 0th players (here TrackedAgent1) private cards to hand history files track_single_player = True # Bounds for randomizing player stack sizes in reset() low_stack_bbs = 50 high_stack_bbs = 200 hand_history_location = 'hands/' invalid_action_penalty = 0 table = Table(active_players, player_names=player_names, track_single_player=track_single_player, stack_low=low_stack_bbs, stack_high=high_stack_bbs, hand_history_location=hand_history_location, invalid_action_penalty=invalid_action_penalty ) table.seed(1) iteration = 1 while True: if iteration % 50 == 0: table.hand_history_enabled = True active_players = np.random.randint(2, 7) table.n_players = active_players obs = table.reset() for agent in agents: agent.reset() acting_player = int(obs[indices.ACTING_PLAYER]) while True: action = agents[acting_player].get_action(obs) obs, reward, done, _ = table.step(action) #print(f"obs={obs} tipo{type(obs)}") #print(f"reward={reward} tipo{type(reward)}") #print(f"done={done} tipo{type(done)}") if done: # Distribute final rewards for i in range(active_players): agents[i].rewards.append(reward[i]) break else: # This step can be skipped unless invalid action penalty is enabled, # since we only get a reward when the pot is distributed, and the done flag is set agents[acting_player].rewards.append(reward[acting_player]) acting_player = int(obs[indices.ACTING_PLAYER]) iteration += 1 table.hand_history_enabled = False
解决方案
1. 修复pokerenv源码中的卡牌存储逻辑
找到pokerenv/table.py中的_street_transition方法,修改卡牌追加代码:
原代码:
new = self.deck.draw(1) self.cards.append(new)
修改为:
new = self.deck.draw(1)[0] # 取出列表中的单个卡牌整数 self.cards.append(new)
这样self.cards中的元素都是int类型,后续调用Card.int_to_str时就不会触发类型错误。
2. 避免重复调用table.step(action)
注意到示例代码中曾有重复调用table.step(action)的代码(已注释),这会导致环境状态被推进两次,可能引发异常行为。确保每个action只调用一次step方法,保持环境状态的一致性。
内容的提问来源于stack exchange,提问作者darth momin
相关产品推荐
相关产品推荐

