You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用MLflow+SQLite后端时遭遇数据库锁定错误求助

MLflow + SQLite 在RHEL7.9上触发 sqlite3.OperationalError: database is locked 错误(新创建数据库)

问题场景

我编写了用于MLflow实验追踪的脚本test.py:

import mlflow
import random
mlflow.set_tracking_uri("sqlite:///mlflow.db")
experiment_id = mlflow.set_experiment('test_experiment')

with mlflow.start_run() as run:
    for i in range(10):
        mlflow.log_metric(key='mse', value=i*random.random(), step=i)

在Red Hat Enterprise Linux Server 7.9上执行时,触发错误:

(sqlite3.OperationalError) database is locked

错误发生在创建experiments表阶段,且确认执行脚本时mlflow.db不存在,由脚本新建。相同脚本在Ubuntu和Windows环境下运行正常,环境配置为Python3.8.11、MLflow1.30.0、SQLAlchemy1.4.46。


可能的解决方法

1. 强制SQLite使用WAL日志模式

RHEL7.9的文件系统锁机制与其他系统存在差异,WAL模式能有效降低锁冲突概率。修改脚本,手动初始化数据库并设置日志模式:

import mlflow
import random
from sqlalchemy import create_engine

mlflow.set_tracking_uri("sqlite:///mlflow.db")

# 初始化数据库并启用WAL模式
engine = create_engine("sqlite:///mlflow.db")
with engine.connect() as conn:
    conn.execute("PRAGMA journal_mode=WAL;")

experiment_id = mlflow.set_experiment('test_experiment')

with mlflow.start_run() as run:
    for i in range(10):
        mlflow.log_metric(key='mse', value=i*random.random(), step=i)

2. 检查SELinux与文件权限

RHEL7.9默认启用的SELinux可能限制SQLite的文件操作,即使文件是新建的也会触发锁异常:

  • 临时关闭SELinux测试:执行sudo setenforce 0后重新运行脚本。若问题解决,需调整SELinux策略,例如为脚本所在目录添加规则:
    sudo semanage fcontext -a -t httpd_sys_rw_content_t "/path/to/your/script/dir(/.*)?"
    sudo restorecon -Rv /path/to/your/script/dir
    
  • 确认目录权限:确保当前用户对脚本所在目录有完整读写权限,执行chmod u+rwx ./后重试。

3. 升级SQLite版本

RHEL7.9默认的SQLite版本(通常为3.7.x)较旧,无法适配MLflow和SQLAlchemy的新特性。手动升级到3.24+版本:

# 下载源码包
wget https://www.sqlite.org/2023/sqlite-autoconf-3410200.tar.gz
# 编译安装
tar xzf sqlite-autoconf-3410200.tar.gz
cd sqlite-autoconf-3410200
./configure --prefix=/usr/local
make && sudo make install
# 更新系统库链接
sudo ldconfig

完成后重新安装Python的sqlite3模块,确保Python调用新安装的SQLite库。

4. 延迟实验运行初始化

MLflow设置实验时可能存在隐式并行操作,短暂等待数据库表创建完成后再启动运行:

import mlflow
import random
import time

mlflow.set_tracking_uri("sqlite:///mlflow.db")
experiment_id = mlflow.set_experiment('test_experiment')

# 等待数据库表初始化完成
time.sleep(1)

with mlflow.start_run() as run:
    for i in range(10):
        mlflow.log_metric(key='mse', value=i*random.random(), step=i)

内容的提问来源于stack exchange,提问作者user17788510

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 02:02:01