You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure ML中部署Databricks模型为评分Web服务失败求助

Azure Databricks集成Azure ML部署模型失败问题排查与解决

问题背景

在Azure Databricks实验中,完成模型训练后,尝试将模型部署到Azure Machine Learning的Azure Container Instance(ACI)时出现部署失败,容器应用崩溃,报错指向评分脚本score.py的init()函数。

报错信息

"Service deployment polling reached non-successful terminal state, current service state: Failed 
code": "AciDeploymentFailed",
  "statusCode": 400,
  "message": "Aci Deployment failed with exception: Your container application crashed. This may be caused by errors in your scoring file's init() function.

关联代码

部署代码

aci_service_name='nyc-taxi-service'

service = Model.deploy(workspace=ws,
                       name=aci_service_name,
                       models=[registered_model],
                       inference_config=inference_config,
                       deployment_config= aci_config, 
                       overwrite=True)

service.wait_for_deployment(show_output=True)
print(service.state)

评分脚本(score.py)

import json
import numpy as np
import pandas as pd
import sklearn
import joblib
from azureml.core.model import Model

columns = ['passengerCount', 'tripDistance', 'hour_of_day', 'day_of_week', 
           'month_num', 'normalizeHolidayName', 'isPaidTimeOff', 'snowDepth', 
           'precipTime', 'precipDepth', 'temperature']

def init():
    global model
    model_path = Model.get_model_path('nyc-taxi-fare')
    model = joblib.load(model_path)
    print('model loaded')

def run(input_json):
    inputs = json.loads(input_json)
    data_df = pd.DataFrame(np.array(inputs).reshape(-1, len(columns)), columns = columns)
    predictions = model.predict(data_df)
    return {'predictions': predictions.tolist()}

排查与修复方案

1. 验证模型路径正确性

  • 确认Azure ML中注册的模型名称与Model.get_model_path传入的名称完全一致(包括版本号,若注册时指定了版本)。若模型带版本,需修改路径获取逻辑:
    # 示例:指定版本号获取模型路径
    model_path = Model.get_model_path(model_name='nyc-taxi-fare', version=1)
    
  • 在init()函数中添加路径打印,便于排查实际路径是否存在:
    print(f"Resolved model path: {model_path}")
    

2. 对齐依赖包版本

容器崩溃常见原因是依赖包版本不匹配(训练环境与推理环境的包版本差异)。需确保inference_config中的conda环境包含与训练环境一致的包版本,示例配置:

name: taxi-inference-env
dependencies:
- python=3.8
- pip:
  - azureml-defaults>=1.42.0
  - pandas==1.3.5
  - numpy==1.21.6
  - scikit-learn==1.0.2
  - joblib==1.1.0

3. 增强init()函数的异常捕获

原代码未处理异常,无法定位具体错误。修改init()函数添加异常捕获与日志输出:

def init():
    global model
    try:
        model_path = Model.get_model_path('nyc-taxi-fare')
        print(f"Resolved model path: {model_path}")
        model = joblib.load(model_path)
        print('Model loaded successfully')
    except Exception as e:
        print(f"Error during model initialization: {str(e)}")
        raise e  # 抛出异常便于Azure ML捕获详细日志

4. 提前验证模型可用性

在Databricks环境中直接测试模型加载逻辑,确认模型文件无损坏:

import joblib
import pandas as pd

# 加载DBFS中的模型文件测试
model = joblib.load('/dbfs/path/to/your/trained/model.joblib')
# 用测试数据验证预测功能
columns = ['passengerCount', 'tripDistance', 'hour_of_day', 'day_of_week', 
           'month_num', 'normalizeHolidayName', 'isPaidTimeOff', 'snowDepth', 
           'precipTime', 'precipDepth', 'temperature']
test_data = pd.DataFrame([[1, 2.5, 10, 3, 10, False, False, 0, 0, 0, 15]], columns=columns)
print(model.predict(test_data))

验证步骤

  1. 应用上述修改后,重新执行部署代码。
  2. 查看service.state,若显示Healthy则部署成功。
  3. 发送测试请求验证预测功能:
import json

test_input = json.dumps([[1, 2.5, 10, 3, 10, False, False, 0, 0, 0, 15]])
response = service.run(input_data=test_input)
print(response)

内容的提问来源于stack exchange,提问作者Michiel Voortman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 09:40:32