训练神经网络玩石头剪刀布时陷入循环的技术求助
训练神经网络击败Abbey石头剪刀布机器人
背景
我正在为freecodecamp的编程课程任务训练神经网络,与4个采用不同预设策略的石头剪刀布机器人对战。作为神经网络自学新手,难以打造出能击败最强机器人Abbey的模型,参考过课程内容及Svilen Todorov的博客。
问题描述
目标是训练模型击败Abbey机器人,其代码如下:
def abbey(prev_opponent_play, opponent_history=[], play_order=[{ "RR": 0, "RP": 0, "RS": 0, "PR": 0, "PP": 0, "PS": 0, "SR": 0, "SP": 0, "SS": 0, }]): if not prev_opponent_play: prev_opponent_play = 'R' opponent_history.append(prev_opponent_play) last_two = "".join(opponent_history[-2:]) if len(last_two) == 2: play_order[0][last_two] += 1 potential_plays = [ prev_opponent_play + "R", prev_opponent_play + "P", prev_opponent_play + "S", ] sub_order = { k: play_order[0][k] for k in potential_plays if k in play_order[0] } prediction = max(sub_order, key=sub_order.get)[-1:] ideal_response = {'P': 'S', 'R': 'P', 'S': 'R'} return ideal_response[prediction]
该机器人为确定性模型,无随机逻辑,会根据对手历史出拳调整策略。
已尝试的解决方案
- 用随机出拳与Abbey对战生成的2000条数据训练模型,10轮epoch后准确率无法突破80%;
- 尝试用当前神经网络与Abbey对战的数据迭代训练模型,但每次迭代后,神经网络很快陷入单步(如R-R-R循环)或多步(如P-S-R循环)出拳循环,被Abbey击败。
现有神经网络代码
import tensorflow as tf from tensorflow import keras from dataManipulation import open_csv import numpy as np import csv TRAINING_ITER = 0 # If at 0, check the file_path manually UPDATING = True RESET_STATES = True NUM_GAMES = 20000 SEQ_LENGTH = 10 # I can feel this is going to be tricky. We are back to len 10. TRAINING = True DENSE_NEURONS = 24 RNN_UNITS = 512 EPOCH_NUM = 2 if TRAINING_ITER == 0: file_path = "db_random_n20000.csv" else: file_path = f'./databases/abbey/db1_nn_abbey_n{NUM_GAMES}_0{TRAINING_ITER-1}.csv' # Raw data load_chk_path = f'./checkpoints/abbey/trained{TRAINING_ITER-1}' last_saved_chk_path = f'./checkpoints/abbey/trained{TRAINING_ITER}' # OHE the data move_to_ohe = { # Create a dictionary to OHE the data "R": [1, 0, 0], "P": [0, 1, 0], "S": [0, 0, 1], } def build_model(seq_l, vocab_size, rnn_units, dense_neurons): this_model = tf.keras.Sequential([ tf.keras.layers.Input(shape=(seq_l, vocab_size), dtype='int32'), tf.keras.layers.Dense(dense_neurons, activation='relu'), tf.keras.layers.LSTM(rnn_units, return_sequences=True), tf.keras.layers.Dense(3, activation='softmax') ]) return this_model if __name__ == "__main__": # Import the data raw_data = open_csv(file_path) data = np.array([[move_to_ohe[val] for val in row] for row in raw_data]) # Next, we need to sequence the data, since the model won't operate on single-move games, nor on games like Todorov's total_seqs = len(data) - (SEQ_LENGTH + 1) # Creating the training sequences Xy1_sequences = np.empty((total_seqs, SEQ_LENGTH + 1, 2, 3), dtype=int) for i, d in enumerate(data): if i == total_seqs: break Xy1_sequences[i] = data[i: i + (SEQ_LENGTH + 1)] x1_seq = Xy1_sequences[:, :-1, :, :] y1_seq = Xy1_sequences[:, 1:, 0, :] # y1 has to be only about the opponent x1_seq, y1_seq = np.reshape(x1_seq, (total_seqs, SEQ_LENGTH, 6)), \ np.reshape(y1_seq, (total_seqs, SEQ_LENGTH, 3)) # I'll try building the model following Russica, since Todorov's method is not working. model = build_model(SEQ_LENGTH, 6, RNN_UNITS, DENSE_NEURONS) print("Database len: ", len(x1_seq)) if TRAINING: model.summary() opt = keras.optimizers.Adam(learning_rate=0.001) if UPDATING: if TRAINING_ITER != 0: model.load_weights(load_chk_path).expect_partial() model.compile(loss='categorical_crossentropy', optimizer=opt, metrics=['accuracy']) model.fit(x1_seq, y1_seq, epochs=EPOCH_NUM, shuffle=False, batch_size=8, verbose=2) # Save the weights model.save_weights(last_saved_chk_path) print("Iteration number ", TRAINING_ITER) print("Weights saved to ", last_saved_chk_path) else: model.load_weights(load_chk_path).expect_partial() # We are OK with working with weights only. This silences warnings. model.summary()
该代码包含与机器人对战的响应逻辑,Play函数为课程任务内置。代码曾成功击败以下简单的Quincy机器人:
def quincy(prev_play, counter=[0]): counter[0] += 1 choices = ["R", "R", "P", "P", "S"] return choices[counter[0] % len(choices)]
因此代码本身应该无bug。我已尝试调整epoch数、激活函数、神经元数量、模型更新方式、LSTM状态、模型重置状态等超参数,但问题仍未解决,寻求模型优化或问题排查的技术建议。
内容的提问来源于stack exchange,提问作者Baldomero
相关产品推荐
相关产品推荐

