You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练自对弈国际象棋神经网络遇train_function空日志错误求助

修复TensorFlow训练国际象棋自对弈模型时的ValueError: Unexpected result of train_function (Empty logs)错误

问题背景

尝试构建通过自对弈学习的国际象棋神经网络,基于TensorFlow编写了包含网络架构、数据集准备、训练及自对弈逻辑的代码,但运行时触发ValueError: Unexpected result of train_function (Empty logs)错误,重写代码、咨询GPT-4后仍未解决,需修复该问题。

错误核心原因

  1. 训练数据集为空:代码中先定义空的dataset数组,直接基于它生成inputs和outputs,导致训练样本数为0,模型无数据可训练,触发空日志错误。
  2. 输入表示逻辑错误:原代码用前一步的走法作为模型输入,而非当前棋盘的完整状态,不符合国际象棋AI的训练逻辑,模型无法学习到棋盘局势与最优走法的关联。
  3. 自对弈与训练的顺序颠倒:原代码先尝试训练模型,再生成自对弈数据,导致训练阶段无有效数据。
  4. 非法走法处理不完善:模型生成的走法未做合法性校验就直接推送,会导致对局中断。

修复方案

1. 调整流程顺序:先生成自对弈初始数据集

先让未训练的模型(或随机策略)进行自对弈,生成初始对局数据,再用这些数据训练模型。

2. 修正输入特征表示

将棋盘状态转换为结构化特征:用13个通道(6种白棋+6种黑棋+空棋盘),每个通道是8×8的矩阵标记对应位置的棋子情况,最后展平为一维数组作为模型输入。

3. 修复训练数据构建逻辑

从自对弈的每一步中提取当前棋盘状态作为输入,下一步的合法走法作为输出标签(用one-hot编码)。

4. 完善非法走法处理

模型预测后,只从合法走法中选择概率最高的走法,避免推送非法走法。

5. 正确启用eager执行

在model.compile()中设置run_eagerly=True,替代全局的tf.config.run_functions_eagerly(True),更精准地排查训练中的问题。

完整修正代码

import tensorflow as tf
import numpy as np
import chess

# 棋盘状态转特征:13个通道(6白棋+6黑棋+空棋盘),展平为一维数组
def board_to_features(board):
    features = np.zeros((13, 8, 8), dtype=np.float32)
    piece_types = [chess.PAWN, chess.KNIGHT, chess.BISHOP, chess.ROOK, chess.QUEEN, chess.KING]
    
    for square in chess.SQUARES:
        piece = board.piece_at(square)
        if piece is None:
            features[12, square // 8, square % 8] = 1.0
        else:
            idx = piece_types.index(piece.piece_type)
            if piece.color == chess.WHITE:
                features[idx, square // 8, square % 8] = 1.0
            else:
                features[idx + 6, square // 8, square % 8] = 1.0
    return features.flatten()

# 走法转索引(64*64=4096种可能)
def move_to_index(move):
    return move.from_square * 64 + move.to_square

# 索引转走法
def index_to_move(index):
    from_square = index // 64
    to_square = index % 64
    return chess.Move(from_square, to_square)

# 定义神经网络架构
model = tf.keras.Sequential([
    tf.keras.layers.Dense(512, activation='relu', input_shape=(832,)),
    tf.keras.layers.Dense(512, activation='relu'),
    tf.keras.layers.Dense(4096, activation='softmax')
])

# 生成自对弈初始数据集
dataset = []
num_games = 100  # 可根据需求调整对局数量

for _ in range(num_games):
    board = chess.Board()
    game_moves = []
    while not board.is_game_over():
        current_features = board_to_features(board)
        legal_moves = list(board.legal_moves)
        if not legal_moves:
            break
        # 初始阶段用随机策略生成走法
        selected_move = np.random.choice(legal_moves)
        game_moves.append((current_features, selected_move))
        board.push(selected_move)
    dataset.extend(game_moves)

# 准备训练数据
X = np.array([item[0] for item in dataset])
y = np.zeros((len(dataset), 4096), dtype=np.float32)
for i, (_, move) in enumerate(dataset):
    y[i, move_to_index(move)] = 1.0

# 训练模型
model.compile(optimizer='adam', loss='categorical_crossentropy', run_eagerly=True)
model.fit(X, y, epochs=5, batch_size=32)

# 训练后自对弈
board = chess.Board()
while not board.is_game_over():
    current_features = board_to_features(board)
    pred_probs = model.predict(np.array([current_features]), verbose=0)[0]
    legal_moves = list(board.legal_moves)
    if not legal_moves:
        break
    # 只保留合法走法的概率
    legal_indices = [move_to_index(move) for move in legal_moves]
    masked_probs = np.zeros_like(pred_probs)
    masked_probs[legal_indices] = pred_probs[legal_indices]
    # 选概率最高的合法走法
    best_move_idx = np.argmax(masked_probs)
    best_move = index_to_move(best_move_idx)
    board.push(best_move)
    print("当前棋盘状态:")
    print(board)
    print("-" * 50)

验证说明

运行修正后的代码:

  1. 模型会先生成初始自对弈数据,再进行正常训练,不会再触发空日志错误。
  2. 训练后的模型会基于棋盘状态选择合法走法进行自对弈,对局可正常进行。

内容的提问来源于stack exchange,提问作者Kavya Sahai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 00:40:11