训练自对弈国际象棋神经网络遇train_function空日志错误求助
修复TensorFlow训练国际象棋自对弈模型时的
ValueError: Unexpected result of train_function (Empty logs)错误 问题背景
尝试构建通过自对弈学习的国际象棋神经网络,基于TensorFlow编写了包含网络架构、数据集准备、训练及自对弈逻辑的代码,但运行时触发ValueError: Unexpected result of train_function (Empty logs)错误,重写代码、咨询GPT-4后仍未解决,需修复该问题。
错误核心原因
- 训练数据集为空:代码中先定义空的
dataset数组,直接基于它生成inputs和outputs,导致训练样本数为0,模型无数据可训练,触发空日志错误。 - 输入表示逻辑错误:原代码用前一步的走法作为模型输入,而非当前棋盘的完整状态,不符合国际象棋AI的训练逻辑,模型无法学习到棋盘局势与最优走法的关联。
- 自对弈与训练的顺序颠倒:原代码先尝试训练模型,再生成自对弈数据,导致训练阶段无有效数据。
- 非法走法处理不完善:模型生成的走法未做合法性校验就直接推送,会导致对局中断。
修复方案
1. 调整流程顺序:先生成自对弈初始数据集
先让未训练的模型(或随机策略)进行自对弈,生成初始对局数据,再用这些数据训练模型。
2. 修正输入特征表示
将棋盘状态转换为结构化特征:用13个通道(6种白棋+6种黑棋+空棋盘),每个通道是8×8的矩阵标记对应位置的棋子情况,最后展平为一维数组作为模型输入。
3. 修复训练数据构建逻辑
从自对弈的每一步中提取当前棋盘状态作为输入,下一步的合法走法作为输出标签(用one-hot编码)。
4. 完善非法走法处理
模型预测后,只从合法走法中选择概率最高的走法,避免推送非法走法。
5. 正确启用eager执行
在model.compile()中设置run_eagerly=True,替代全局的tf.config.run_functions_eagerly(True),更精准地排查训练中的问题。
完整修正代码
import tensorflow as tf import numpy as np import chess # 棋盘状态转特征:13个通道(6白棋+6黑棋+空棋盘),展平为一维数组 def board_to_features(board): features = np.zeros((13, 8, 8), dtype=np.float32) piece_types = [chess.PAWN, chess.KNIGHT, chess.BISHOP, chess.ROOK, chess.QUEEN, chess.KING] for square in chess.SQUARES: piece = board.piece_at(square) if piece is None: features[12, square // 8, square % 8] = 1.0 else: idx = piece_types.index(piece.piece_type) if piece.color == chess.WHITE: features[idx, square // 8, square % 8] = 1.0 else: features[idx + 6, square // 8, square % 8] = 1.0 return features.flatten() # 走法转索引(64*64=4096种可能) def move_to_index(move): return move.from_square * 64 + move.to_square # 索引转走法 def index_to_move(index): from_square = index // 64 to_square = index % 64 return chess.Move(from_square, to_square) # 定义神经网络架构 model = tf.keras.Sequential([ tf.keras.layers.Dense(512, activation='relu', input_shape=(832,)), tf.keras.layers.Dense(512, activation='relu'), tf.keras.layers.Dense(4096, activation='softmax') ]) # 生成自对弈初始数据集 dataset = [] num_games = 100 # 可根据需求调整对局数量 for _ in range(num_games): board = chess.Board() game_moves = [] while not board.is_game_over(): current_features = board_to_features(board) legal_moves = list(board.legal_moves) if not legal_moves: break # 初始阶段用随机策略生成走法 selected_move = np.random.choice(legal_moves) game_moves.append((current_features, selected_move)) board.push(selected_move) dataset.extend(game_moves) # 准备训练数据 X = np.array([item[0] for item in dataset]) y = np.zeros((len(dataset), 4096), dtype=np.float32) for i, (_, move) in enumerate(dataset): y[i, move_to_index(move)] = 1.0 # 训练模型 model.compile(optimizer='adam', loss='categorical_crossentropy', run_eagerly=True) model.fit(X, y, epochs=5, batch_size=32) # 训练后自对弈 board = chess.Board() while not board.is_game_over(): current_features = board_to_features(board) pred_probs = model.predict(np.array([current_features]), verbose=0)[0] legal_moves = list(board.legal_moves) if not legal_moves: break # 只保留合法走法的概率 legal_indices = [move_to_index(move) for move in legal_moves] masked_probs = np.zeros_like(pred_probs) masked_probs[legal_indices] = pred_probs[legal_indices] # 选概率最高的合法走法 best_move_idx = np.argmax(masked_probs) best_move = index_to_move(best_move_idx) board.push(best_move) print("当前棋盘状态:") print(board) print("-" * 50)
验证说明
运行修正后的代码:
- 模型会先生成初始自对弈数据,再进行正常训练,不会再触发空日志错误。
- 训练后的模型会基于棋盘状态选择合法走法进行自对弈,对局可正常进行。
内容的提问来源于stack exchange,提问作者Kavya Sahai
相关产品推荐
相关产品推荐

