You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

mlflow.tensorflow.log_model报NoCredential错误及boto3依赖疑问求助

问题描述
  • 在Kubernetes集群部署了MLflow追踪服务器,使用mlflow.tensorflow.log_model记录Keras模型时遭遇以下问题:
    1. 初始报错no module named boto3,安装boto3后又出现NoCredential错误
    2. 参数与指标可正常记录,但模型制品无法保存
    3. 未使用AWS环境,疑惑该API为何依赖boto3,寻求替代方案

相关代码

tracking_url = "https://......"
mlflow.set_tracking_uri(tracking_url)
mlflow.set_experiment('test_mlflow')

def create_classifier():
    classifier = tf.keras.Sequential()
    classifier.add(tf.keras.layers.Dense(units = 6, kernel_initializer = 'uniform', activation = 'relu', input_dim = 12))
    classifier.add(tf.keras.layers.Dense(units = 6, kernel_initializer = 'uniform', activation = 'relu'))
    classifier.add(tf.keras.layers.Dense(units = 1, kernel_initializer = 'uniform', activation = 'sigmoid'))
    classifier.compile(optimizer = optimizer, loss = loss, metrics = ['accuracy'])
    return classifier

classifier = create_classifier()

history = classifier.fit(X_train, y_train, batch_size = batch_size, epochs = epochs,verbose = 1)

test_score, test_acc = classifier.evaluate(X_test, y_test,
                            batch_size=batch_size)

tf.keras.models.save_model(classifier, model_save_path)

run_name = "sample-ann-run3"
with mlflow.start_run(run_name=run_name):
    mlflow.log_param("batch_size", batch_size)
    mlflow.log_param("learning_rate", learning_rate)
    mlflow.log_param("epochs", epochs)
    mlflow.log_metric("loss", test_score)
    mlflow.log_metric("accuracy", test_acc)

    mlflow.tensorflow.log_model(model=classifier, registered_model_name="sample-ann-1", artifact_path=model_save_path)

核心错误信息

botocore.exceptions.NoCredentialsError: Unable to locate credentials
解决方案

1. 依赖boto3的原因

MLFlow的TensorFlow模型日志模块默认使用TensorFlow的SavedModel格式,而TensorFlow的文件系统适配模块会自动加载S3相关插件——即使你未使用AWS,只要该插件被加载,就会触发boto3的依赖检查,进而引发凭证错误。

2. 两种可行解决思路

思路一:禁用TensorFlow的S3文件系统插件

在代码开头添加环境变量配置,强制TensorFlow不加载S3相关适配逻辑:

import os
os.environ['TF_DISABLE_S3'] = '1'

此方法可直接阻止TensorFlow初始化S3客户端,避免boto3的凭证校验流程。

思路二:改用MLFlow通用模型日志API

放弃mlflow.tensorflow.log_model,使用mlflow.log_model并指定TensorFlow flavor,同样能完成模型记录与注册:

from mlflow.models.signature import infer_signature

# 生成模型签名(可选,推荐用于明确输入输出格式)
signature = infer_signature(X_train, classifier.predict(X_train))

# 通用API日志模型
mlflow.log_model(
    model=classifier,
    artifact_path=model_save_path,
    registered_model_name="sample-ann-1",
    signature=signature,
    flavor=mlflow.tensorflow
)

该方法绕开TensorFlow内部的文件系统适配逻辑,直接通过MLFlow核心模块处理模型,避免不必要的boto3依赖。

3. 额外注意事项

  • 确认MLFlow追踪服务器的制品存储路径配置为集群内存储(如PVC)或本地存储,而非S3类存储,否则服务器端仍可能出现存储适配问题。
  • 若必须保留mlflow.tensorflow.log_model,可临时配置空AWS凭证跳过检查(仅作临时 workaround,不推荐长期使用):
    import boto3
    boto3.client('s3', aws_access_key_id='dummy', aws_secret_access_key='dummy')
    

内容的提问来源于stack exchange,提问作者Henry Bai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 04:05:25