Python加载PGN棋谱时内存占用过高的原因排查
解决PGN文件加载时Python进程内存占用过高问题
问题原因
__sizeof__的局限性:你调用的games.__sizeof__()仅计算了列表容器本身的内存开销,完全没有包含列表中每个chess.pgn.Game对象的实际内存占用。每个Game对象会存储完整的对局树、每一步的棋盘状态快照、元数据(选手信息、时间控制等),这些数据的内存总和远大于列表本身的大小。- chess库的解析特性:
chess.pgn.read_game默认会构建完整的游戏对象,包含所有走法对应的棋盘状态,这类数据会随着对局数量增加快速累积内存。
解决方案
1. 只提取必要信息,丢弃完整Game对象
不要存储整个Game对象,而是提取你需要的核心数据(比如走法序列、对局结果、选手等级分等),用轻量结构(如字典)存储,大幅降低内存占用。
示例代码:
import chess.pgn from tqdm import tqdm def load_game_metadata(n_games: int) -> list[dict]: """仅提取必要的对局元数据和走法,不保存完整Game对象""" with open("files\\lichess_elite_2022-04.pgn") as pgn_file: game_data = [] for i in tqdm(range(n_games), desc="Loading games", unit=" games"): game = chess.pgn.read_game(pgn_file) if game is None: break # 按需提取所需字段 data = { "white": game.headers.get("White"), "white_elo": game.headers.get("WhiteElo"), "black": game.headers.get("Black"), "black_elo": game.headers.get("BlackElo"), "result": game.headers.get("Result"), "moves": [move.uci() for move in game.mainline_moves()] } game_data.append(data) # 手动释放Game对象引用,辅助垃圾回收 del game return game_data game_metadata = load_game_metadata(10000)
2. 使用Visitor模式解析PGN(内存最优)
chess库支持Visitor模式,无需构建完整的Game对象,直接在解析过程中提取数据,内存占用最低。这种方式不会创建任何冗余的Game对象,完全按需收集信息。
示例代码:
import chess.pgn from tqdm import tqdm class GameDataVisitor(chess.pgn.BaseVisitor): def __init__(self): self.current_game = {} self.moves = [] self.all_games = [] def visit_header(self, tagname, tagvalue): # 收集表头元数据 self.current_game[tagname] = tagvalue def visit_move(self, board, move): # 收集主线路走法 self.moves.append(move.uci()) def end_game(self): # 一局解析完成后整合数据并重置状态 self.current_game["moves"] = self.moves self.all_games.append(self.current_game) self.current_game = {} self.moves = [] def load_games_with_visitor(n_games: int) -> list[dict]: visitor = GameDataVisitor() with open("files\\lichess_elite_2022-04.pgn") as pgn_file: for _ in tqdm(range(n_games), desc="Loading games", unit=" games"): # 通过visitor解析,不生成完整Game对象 if chess.pgn.read_game(pgn_file, visitor=visitor) is None: break return visitor.all_games game_data = load_games_with_visitor(10000)
3. 分批处理并写入文件
如果需要处理20万局这类大规模数据,不要一次性将所有数据保存在内存中,而是处理一批就写入文件(比如CSV、JSON Lines格式),然后清空内存继续处理下一批。
示例代码:
import chess.pgn import csv from tqdm import tqdm def process_batch(pgn_file, batch_size: int, writer): for _ in range(batch_size): game = chess.pgn.read_game(pgn_file) if game is None: return False # 提取数据并写入 row = { "White": game.headers.get("White"), "WhiteElo": game.headers.get("WhiteElo"), "Black": game.headers.get("Black"), "BlackElo": game.headers.get("BlackElo"), "Result": game.headers.get("Result"), "Moves": " ".join(m.uci() for m in game.mainline_moves()) } writer.writerow(row) del game return True def process_large_pgn(total_games: int, batch_size: int = 1000): with open("files\\lichess_elite_2022-04.pgn") as pgn_file, \ open("game_data.csv", "w", newline="", encoding="utf-8") as csv_file: fieldnames = ["White", "WhiteElo", "Black", "BlackElo", "Result", "Moves"] writer = csv.DictWriter(csv_file, fieldnames=fieldnames) writer.writeheader() processed = 0 with tqdm(total=total_games, desc="Processing games", unit=" games") as pbar: while processed < total_games: current_batch = min(batch_size, total_games - processed) if not process_batch(pgn_file, current_batch, writer): break processed += current_batch pbar.update(current_batch) # 处理20万局对局 process_large_pgn(200000)
内容的提问来源于stack exchange,提问作者Sandro Martens
相关产品推荐
相关产品推荐

