You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python加载PGN棋谱时内存占用过高的原因排查

解决PGN文件加载时Python进程内存占用过高问题

问题原因

  1. __sizeof__的局限性:你调用的games.__sizeof__()仅计算了列表容器本身的内存开销,完全没有包含列表中每个chess.pgn.Game对象的实际内存占用。每个Game对象会存储完整的对局树、每一步的棋盘状态快照、元数据(选手信息、时间控制等),这些数据的内存总和远大于列表本身的大小。
  2. chess库的解析特性:chess.pgn.read_game默认会构建完整的游戏对象,包含所有走法对应的棋盘状态,这类数据会随着对局数量增加快速累积内存。

解决方案

1. 只提取必要信息,丢弃完整Game对象

不要存储整个Game对象,而是提取你需要的核心数据(比如走法序列、对局结果、选手等级分等),用轻量结构(如字典)存储,大幅降低内存占用。

示例代码:

import chess.pgn
from tqdm import tqdm

def load_game_metadata(n_games: int) -> list[dict]:
    """仅提取必要的对局元数据和走法,不保存完整Game对象"""
    with open("files\\lichess_elite_2022-04.pgn") as pgn_file:
        game_data = []
        for i in tqdm(range(n_games), desc="Loading games", unit=" games"):
            game = chess.pgn.read_game(pgn_file)
            if game is None:
                break
            # 按需提取所需字段
            data = {
                "white": game.headers.get("White"),
                "white_elo": game.headers.get("WhiteElo"),
                "black": game.headers.get("Black"),
                "black_elo": game.headers.get("BlackElo"),
                "result": game.headers.get("Result"),
                "moves": [move.uci() for move in game.mainline_moves()]
            }
            game_data.append(data)
            # 手动释放Game对象引用,辅助垃圾回收
            del game
        return game_data

game_metadata = load_game_metadata(10000)

2. 使用Visitor模式解析PGN(内存最优)

chess库支持Visitor模式,无需构建完整的Game对象,直接在解析过程中提取数据,内存占用最低。这种方式不会创建任何冗余的Game对象,完全按需收集信息。

示例代码:

import chess.pgn
from tqdm import tqdm

class GameDataVisitor(chess.pgn.BaseVisitor):
    def __init__(self):
        self.current_game = {}
        self.moves = []
        self.all_games = []

    def visit_header(self, tagname, tagvalue):
        # 收集表头元数据
        self.current_game[tagname] = tagvalue

    def visit_move(self, board, move):
        # 收集主线路走法
        self.moves.append(move.uci())

    def end_game(self):
        # 一局解析完成后整合数据并重置状态
        self.current_game["moves"] = self.moves
        self.all_games.append(self.current_game)
        self.current_game = {}
        self.moves = []

def load_games_with_visitor(n_games: int) -> list[dict]:
    visitor = GameDataVisitor()
    with open("files\\lichess_elite_2022-04.pgn") as pgn_file:
        for _ in tqdm(range(n_games), desc="Loading games", unit=" games"):
            # 通过visitor解析,不生成完整Game对象
            if chess.pgn.read_game(pgn_file, visitor=visitor) is None:
                break
    return visitor.all_games

game_data = load_games_with_visitor(10000)

3. 分批处理并写入文件

如果需要处理20万局这类大规模数据,不要一次性将所有数据保存在内存中,而是处理一批就写入文件(比如CSV、JSON Lines格式),然后清空内存继续处理下一批。

示例代码:

import chess.pgn
import csv
from tqdm import tqdm

def process_batch(pgn_file, batch_size: int, writer):
    for _ in range(batch_size):
        game = chess.pgn.read_game(pgn_file)
        if game is None:
            return False
        # 提取数据并写入
        row = {
            "White": game.headers.get("White"),
            "WhiteElo": game.headers.get("WhiteElo"),
            "Black": game.headers.get("Black"),
            "BlackElo": game.headers.get("BlackElo"),
            "Result": game.headers.get("Result"),
            "Moves": " ".join(m.uci() for m in game.mainline_moves())
        }
        writer.writerow(row)
        del game
    return True

def process_large_pgn(total_games: int, batch_size: int = 1000):
    with open("files\\lichess_elite_2022-04.pgn") as pgn_file, \
         open("game_data.csv", "w", newline="", encoding="utf-8") as csv_file:
        fieldnames = ["White", "WhiteElo", "Black", "BlackElo", "Result", "Moves"]
        writer = csv.DictWriter(csv_file, fieldnames=fieldnames)
        writer.writeheader()

        processed = 0
        with tqdm(total=total_games, desc="Processing games", unit=" games") as pbar:
            while processed < total_games:
                current_batch = min(batch_size, total_games - processed)
                if not process_batch(pgn_file, current_batch, writer):
                    break
                processed += current_batch
                pbar.update(current_batch)

# 处理20万局对局
process_large_pgn(200000)

内容的提问来源于stack exchange,提问作者Sandro Martens

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 13:45:28