You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Slurm并行调参时JSON解码随机报错,本地无法复现求助

解决Slurm并行任务中JSON日志文件随机解码错误问题

问题背景

在Compute Canada通过Slurm提交32个并行任务调优Sklearn/TensorFlow模型时,随机出现json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)错误,本地单任务运行无此问题,日志文件事后查看内容正常,将后缀从.json改为.txt仅略有好转。

根因分析

核心问题是并发读写的竞态条件:多个任务同时对同一个日志文件执行读、写操作,当一个任务正在写入文件(比如json.dump未完成)时,另一个任务开始读取,此时文件内容是不完整的JSON结构,触发解码错误。修改文件后缀只是巧合,无法解决本质问题。

解决方案

在文件读写操作前后添加文件锁,确保同一时间只有一个进程能访问日志文件。Linux环境下可以用fcntl模块实现文件锁。

修改后的代码示例

import os
import json
import fcntl
from typing import Union, List, Callable
import numpy as np
import pandas as pd
import optuna

def optunize_hyperparameters(X_tr: Union[List[np.ndarray], pd.DataFrame, np.ndarray], y_tr: Union[dict, pd.DataFrame, np.ndarray],
                            Objective: Callable, builder_func: Callable, model_name: str, fit_params: dict, log_path: str, n_trials: int=10, **kwargs) -> dict:
    """
    采用网格搜索优化模型超参数,若存在已保存的超参数则加载。

    参数
    ----------
    X_tr: 
        训练特征。
    y_tr: 
        训练目标
    Objective:
        Optuna目标可调用类。
    model_name:
        格式为"<模型名称>_qx",其中x取值为{5,25,50,75,95}。
    fit_params:
        传递给model.fit()的参数
    log_path:
        超参数日志文件路径,例如"tuning_log.txt"

    返回值
    -------
    best_hps:
        可作为**kwargs传入builder_func的字典。
    """
    
    # 确保日志文件存在,带锁操作
    if not os.path.exists(log_path):
        with open(log_path, 'w') as f:
            # 加排他锁
            fcntl.flock(f, fcntl.LOCK_EX)
            json.dump({}, f)
            fcntl.flock(f, fcntl.LOCK_UN)

    log = {}
    # 带锁加载日志
    try:
        with open(log_path, 'r+') as f:
            fcntl.flock(f, fcntl.LOCK_SH)  # 共享读锁
            log = json.load(f)
            fcntl.flock(f, fcntl.LOCK_UN)
            print("Successfully loaded existing hyperparameters.")
    except OSError as e: 
        print(e)
    except json.JSONDecodeError as e:
        print(f"JSON decode error, retrying...: {e}")
        # 重试一次,避免锁释放不完全导致的读取问题
        with open(log_path, 'r+') as f:
            fcntl.flock(f, fcntl.LOCK_SH)
            log = json.load(f)
            fcntl.flock(f, fcntl.LOCK_UN)

    # 查找现有超参数,否则进行优化
    try:
        best_hps = log[model_name]
        print("Existing hyperparameters found, loading...")
    
    except KeyError:
        print("No existing hyperparameters found, optimizing hyperparameters...")
        study = optuna.create_study(
            sampler=optuna.samplers.RandomSampler(),
            pruner=optuna.pruners.SuccessiveHalvingPruner(),
            direction='maximize'
        )

        study.optimize(
            Objective(
                X_tr, y_tr,
                builder_func=builder_func,
                fit_params=fit_params,
                **kwargs
            ),
            n_trials=n_trials,
            n_jobs=-1
        )
        
        best_hps = study.best_params
        
        # 将超参数加入日志并保存,带排他锁
        with open(log_path, 'r+') as f:
            fcntl.flock(f, fcntl.LOCK_EX)  # 排他写锁
            # 重新加载最新日志,避免其他任务已修改
            f.seek(0)
            log = json.load(f)
            log[model_name] = best_hps
            f.seek(0)
            f.truncate()  # 清空文件后写入
            json.dump(log, f)
            fcntl.flock(f, fcntl.LOCK_UN)
        
    return best_hps

关键修改点

  • 文件锁机制:使用fcntl.flock实现共享读锁(LOCK_SH)和排他写锁(LOCK_EX),确保读写操作互斥。
  • 写操作前重新读日志:在写入新超参数前重新加载日志,避免覆盖其他任务刚写入的内容。
  • 异常重试:针对JSON解码错误添加重试逻辑,处理锁释放不及时的边缘情况。

备选方案

如果文件锁在集群环境下仍有问题,可以尝试:

  • 为每个任务分配独立的日志文件(比如log_{model_name}.txt),训练完成后再合并所有日志。
  • 使用分布式键值存储代替本地文件存储超参数,避免文件系统的并发问题。

内容的提问来源于stack exchange,提问作者BigBossRob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 05:10:04