You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练神经网络玩石头剪刀布时陷入循环的技术求助

训练神经网络击败Abbey石头剪刀布机器人

背景

我正在为freecodecamp的编程课程任务训练神经网络,与4个采用不同预设策略的石头剪刀布机器人对战。作为神经网络自学新手,难以打造出能击败最强机器人Abbey的模型,参考过课程内容及Svilen Todorov的博客。

问题描述

目标是训练模型击败Abbey机器人,其代码如下:

def abbey(prev_opponent_play,
          opponent_history=[],
          play_order=[{
              "RR": 0,
              "RP": 0,
              "RS": 0,
              "PR": 0,
              "PP": 0,
              "PS": 0,
              "SR": 0,
              "SP": 0,
              "SS": 0,
          }]):

    if not prev_opponent_play:
        prev_opponent_play = 'R'
    opponent_history.append(prev_opponent_play)

    last_two = "".join(opponent_history[-2:])
    if len(last_two) == 2:
        play_order[0][last_two] += 1

    potential_plays = [
        prev_opponent_play + "R",
        prev_opponent_play + "P",
        prev_opponent_play + "S",
    ]

    sub_order = {
        k: play_order[0][k]
        for k in potential_plays if k in play_order[0]
    }

    prediction = max(sub_order, key=sub_order.get)[-1:]

    ideal_response = {'P': 'S', 'R': 'P', 'S': 'R'}
    return ideal_response[prediction]

该机器人为确定性模型,无随机逻辑,会根据对手历史出拳调整策略。

已尝试的解决方案

  • 用随机出拳与Abbey对战生成的2000条数据训练模型,10轮epoch后准确率无法突破80%;
  • 尝试用当前神经网络与Abbey对战的数据迭代训练模型,但每次迭代后,神经网络很快陷入单步(如R-R-R循环)或多步(如P-S-R循环)出拳循环,被Abbey击败。

现有神经网络代码

import tensorflow as tf
from tensorflow import keras
from dataManipulation import open_csv
import numpy as np
import csv

TRAINING_ITER = 0   # If at 0, check the file_path manually

UPDATING = True
RESET_STATES = True

NUM_GAMES = 20000
SEQ_LENGTH = 10  # I can feel this is going to be tricky. We are back to len 10.
TRAINING = True
DENSE_NEURONS = 24
RNN_UNITS = 512
EPOCH_NUM = 2


if TRAINING_ITER == 0:
    file_path = "db_random_n20000.csv"
else:
    file_path = f'./databases/abbey/db1_nn_abbey_n{NUM_GAMES}_0{TRAINING_ITER-1}.csv'  # Raw data
load_chk_path = f'./checkpoints/abbey/trained{TRAINING_ITER-1}'
last_saved_chk_path = f'./checkpoints/abbey/trained{TRAINING_ITER}'

# OHE the data
move_to_ohe = {  # Create a dictionary to OHE the data
    "R": [1, 0, 0],
    "P": [0, 1, 0],
    "S": [0, 0, 1],
}


def build_model(seq_l, vocab_size, rnn_units, dense_neurons):
    this_model = tf.keras.Sequential([
        tf.keras.layers.Input(shape=(seq_l, vocab_size), dtype='int32'),
        tf.keras.layers.Dense(dense_neurons, activation='relu'),
        tf.keras.layers.LSTM(rnn_units,
                             return_sequences=True),
        tf.keras.layers.Dense(3, activation='softmax')
    ])
    return this_model


if __name__ == "__main__":
    # Import the data
    raw_data = open_csv(file_path)
    data = np.array([[move_to_ohe[val] for val in row] for row in raw_data])  
    # Next, we need to sequence the data, since the model won't operate on single-move games, nor on games like Todorov's
    total_seqs = len(data) - (SEQ_LENGTH + 1)

    # Creating the training sequences
    Xy1_sequences = np.empty((total_seqs, SEQ_LENGTH + 1, 2, 3), dtype=int)
    for i, d in enumerate(data):
        if i == total_seqs:
            break
        Xy1_sequences[i] = data[i: i + (SEQ_LENGTH + 1)]
    x1_seq = Xy1_sequences[:, :-1, :, :]
    y1_seq = Xy1_sequences[:, 1:, 0, :]  # y1 has to be only about the opponent        

    x1_seq, y1_seq = np.reshape(x1_seq, (total_seqs, SEQ_LENGTH, 6)), \
        np.reshape(y1_seq, (total_seqs, SEQ_LENGTH, 3))


    # I'll try building the model following Russica, since Todorov's method is not working.

    model = build_model(SEQ_LENGTH, 6, RNN_UNITS, DENSE_NEURONS)
    print("Database len: ", len(x1_seq))

    if TRAINING:
        model.summary()
        opt = keras.optimizers.Adam(learning_rate=0.001)
        if UPDATING:
            if TRAINING_ITER != 0:
                model.load_weights(load_chk_path).expect_partial()
        model.compile(loss='categorical_crossentropy', optimizer=opt,
                      metrics=['accuracy'])
        model.fit(x1_seq, y1_seq, epochs=EPOCH_NUM, shuffle=False, batch_size=8, verbose=2)

        # Save the weights
        model.save_weights(last_saved_chk_path)
        print("Iteration number ", TRAINING_ITER)
        print("Weights saved to ", last_saved_chk_path)
    else:
        model.load_weights(load_chk_path).expect_partial()  # We are OK with working with weights only. This silences warnings.
        model.summary()

该代码包含与机器人对战的响应逻辑,Play函数为课程任务内置。代码曾成功击败以下简单的Quincy机器人:

def quincy(prev_play, counter=[0]):
    counter[0] += 1
    choices = ["R", "R", "P", "P", "S"]
    return choices[counter[0] % len(choices)]

因此代码本身应该无bug。我已尝试调整epoch数、激活函数、神经元数量、模型更新方式、LSTM状态、模型重置状态等超参数,但问题仍未解决,寻求模型优化或问题排查的技术建议。


内容的提问来源于stack exchange,提问作者Baldomero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 10:54:57