You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用井字棋对局数据集训练自定义深度强化学习网络?

用井字棋对局数据训练你的模型步骤

1. 把字符棋盘转成数值

神经网络只能处理数值,先给棋盘字符做编码映射:

  • _(空位置)→ 0
  • X(己方棋子)→ 1
  • O(对方棋子)→ -1

实现编码函数:

def encode_board(board):
    char_to_num = {'_': 0, 'X': 1, 'O': -1}
    return [char_to_num[cell] for cell in board]

# 处理原始对局数据
raw_data = [[['_', '_', '_', '_', '_', '_', '_', '_', '_'], ['_', '_', '_', '_', 'X', '_', '_', '_', 'O'], ...]]  # 你的原始数据集
encoded_games = []
for game in raw_data:
    encoded_states = [encode_board(state) for state in game]
    encoded_games.append(encoded_states)

2. 提取训练用的输入-输出对

你的对局是连续状态序列,每个状态对应的下一个动作就是模型要学习的“最优动作”。需要从每局中提取:

  • 输入:当前编码后的棋盘状态
  • 输出:标记该状态下实际走的动作位置(让模型知道这个位置是最优选择)

通过对比相邻状态找到动作位置:

def find_action_idx(current_state, next_state):
    for idx in range(9):
        if current_state[idx] != next_state[idx]:
            return idx
    return None  # 对局结束,无有效动作

# 生成训练集
X_train = []
y_train = []
for game in encoded_games:
    # 遍历每局的连续状态对
    for i in range(len(game)-1):
        curr_state = game[i]
        next_state = game[i+1]
        action_idx = find_action_idx(curr_state, next_state)
        if action_idx is not None:
            X_train.append(curr_state)
            # 构造目标数组:仅动作位置设为1,其余为0(确保模型输出该位置数值最高)
            target = [0]*9
            target[action_idx] = 1
            y_train.append(target)

# 转成numpy数组适配Keras
import numpy as np
X_train = np.array(X_train)
y_train = np.array(y_train)

3. 初始化模型

你的agent函数需要明确输入形状,这里输入是9维一维数组,传入state_shape=(9,):

import tensorflow as tf
from tensorflow import keras

def agent(state_shape, action_shape):
    learning_rate = 0.001
    init = tf.keras.initializers.HeUniform()
    model = keras.Sequential()
    model.add(keras.layers.Dense(24, input_shape=state_shape, activation='relu', kernel_initializer=init))
    model.add(keras.layers.Dense(12, activation='relu', kernel_initializer=init))
    model.add(keras.layers.Dense(action_shape, activation='linear', kernel_initializer=init))
    model.compile(loss=tf.keras.losses.Huber(), optimizer=tf.keras.optimizers.Adam(lr=learning_rate), metrics=['accuracy'])
    return model

# 创建模型实例
model = agent(state_shape=(9,), action_shape=9)

4. 启动训练

用Keras的fit方法执行训练,参数可根据效果调整:

training_history = model.fit(
    X_train, y_train,
    epochs=50,  # 先跑50轮观察效果
    batch_size=8,
    validation_split=0.2  # 用20%数据做验证,检测过拟合
)

5. 测试模型效果

训练完成后,输入棋盘状态即可预测最优动作:

# 测试空棋盘
test_board = ['_', '_', '_', '_', '_', '_', '_', '_', '_']
encoded_test = np.array([encode_board(test_board)])
preds = model.predict(encoded_test)[0]
best_action = np.argmax(preds)  # 取数值最高的索引作为最优动作
print(f"最优动作位置:{best_action}(棋盘从左到右、从上到下数第{best_action+1}个位置)")

额外优化建议

  • 如果对局包含胜负结果,可以给获胜动作设更高目标值(比如2),失败动作设负值(比如-1),提升模型策略能力
  • 数据量不足时,可让模型自我对局生成更多训练数据(DQN强化学习思路),先把监督学习跑通再尝试
  • 若训练准确率偏低,可调整模型神经元数量、学习率或增加训练轮数

内容的提问来源于stack exchange,提问作者Omar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 09:02:50