Azure ML中部署Databricks模型为评分Web服务失败求助
Azure Databricks集成Azure ML部署模型失败问题排查与解决
问题背景
在Azure Databricks实验中,完成模型训练后,尝试将模型部署到Azure Machine Learning的Azure Container Instance(ACI)时出现部署失败,容器应用崩溃,报错指向评分脚本score.py的init()函数。
报错信息
"Service deployment polling reached non-successful terminal state, current service state: Failed code": "AciDeploymentFailed", "statusCode": 400, "message": "Aci Deployment failed with exception: Your container application crashed. This may be caused by errors in your scoring file's init() function.
关联代码
部署代码
aci_service_name='nyc-taxi-service' service = Model.deploy(workspace=ws, name=aci_service_name, models=[registered_model], inference_config=inference_config, deployment_config= aci_config, overwrite=True) service.wait_for_deployment(show_output=True) print(service.state)
评分脚本(score.py)
import json import numpy as np import pandas as pd import sklearn import joblib from azureml.core.model import Model columns = ['passengerCount', 'tripDistance', 'hour_of_day', 'day_of_week', 'month_num', 'normalizeHolidayName', 'isPaidTimeOff', 'snowDepth', 'precipTime', 'precipDepth', 'temperature'] def init(): global model model_path = Model.get_model_path('nyc-taxi-fare') model = joblib.load(model_path) print('model loaded') def run(input_json): inputs = json.loads(input_json) data_df = pd.DataFrame(np.array(inputs).reshape(-1, len(columns)), columns = columns) predictions = model.predict(data_df) return {'predictions': predictions.tolist()}
排查与修复方案
1. 验证模型路径正确性
- 确认Azure ML中注册的模型名称与
Model.get_model_path传入的名称完全一致(包括版本号,若注册时指定了版本)。若模型带版本,需修改路径获取逻辑:# 示例:指定版本号获取模型路径 model_path = Model.get_model_path(model_name='nyc-taxi-fare', version=1) - 在
init()函数中添加路径打印,便于排查实际路径是否存在:print(f"Resolved model path: {model_path}")
2. 对齐依赖包版本
容器崩溃常见原因是依赖包版本不匹配(训练环境与推理环境的包版本差异)。需确保inference_config中的conda环境包含与训练环境一致的包版本,示例配置:
name: taxi-inference-env dependencies: - python=3.8 - pip: - azureml-defaults>=1.42.0 - pandas==1.3.5 - numpy==1.21.6 - scikit-learn==1.0.2 - joblib==1.1.0
3. 增强init()函数的异常捕获
原代码未处理异常,无法定位具体错误。修改init()函数添加异常捕获与日志输出:
def init(): global model try: model_path = Model.get_model_path('nyc-taxi-fare') print(f"Resolved model path: {model_path}") model = joblib.load(model_path) print('Model loaded successfully') except Exception as e: print(f"Error during model initialization: {str(e)}") raise e # 抛出异常便于Azure ML捕获详细日志
4. 提前验证模型可用性
在Databricks环境中直接测试模型加载逻辑,确认模型文件无损坏:
import joblib import pandas as pd # 加载DBFS中的模型文件测试 model = joblib.load('/dbfs/path/to/your/trained/model.joblib') # 用测试数据验证预测功能 columns = ['passengerCount', 'tripDistance', 'hour_of_day', 'day_of_week', 'month_num', 'normalizeHolidayName', 'isPaidTimeOff', 'snowDepth', 'precipTime', 'precipDepth', 'temperature'] test_data = pd.DataFrame([[1, 2.5, 10, 3, 10, False, False, 0, 0, 0, 15]], columns=columns) print(model.predict(test_data))
验证步骤
- 应用上述修改后,重新执行部署代码。
- 查看
service.state,若显示Healthy则部署成功。 - 发送测试请求验证预测功能:
import json test_input = json.dumps([[1, 2.5, 10, 3, 10, False, False, 0, 0, 0, 15]]) response = service.run(input_data=test_input) print(response)
内容的提问来源于stack exchange,提问作者Michiel Voortman
相关产品推荐
相关产品推荐

