You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Data Factory自定义批处理服务运行失败(错误码2500)求助

解决Azure Data Factory运行自定义Python训练脚本时的FailureExitCode错误

问题现象

运行ADF管道中的自定义服务执行train.py脚本时,返回错误:

{"errorCategory":0,"code":"FailureExitCode","message":"The task exited with an exit code representing a failure","details":[{"Name":"Message","Value":"The task process exited with an unexpected exit code"},{"Name":"AdditionalErrorCode","Value":"FailureExitCode"}]}

已完成的配置:

  • 创建存储账户并开启公共/匿名访问
  • 配置ADF与Azure Batch、Blob存储的链接服务
  • 创建带池的Batch账户
  • 参考官方及第三方教程配置,仍无法解决

代码问题排查

从提供的train.py代码来看,存在几个可能导致脚本异常退出的点:

  1. 未处理命令行参数缺失
    在入口块中直接取arguments[1]作为account_key,若ADF传递参数遗漏,会触发IndexError导致脚本退出。

  2. Blob文件查找逻辑缺陷
    初始化blob_target = [],若遍历Blob时找不到traindata/car_data.csv,后续调用SAS生成函数会传递错误参数类型,引发异常。

  3. 无全局异常捕获与日志输出
    脚本未处理运行时错误,任何异常都会直接退出,无法在ADF中查看具体错误信息,仅能得到模糊的退出码提示。

  4. 依赖包未明确声明
    脚本使用pandas、scikit-learn等第三方库,若Batch池环境未预先安装这些依赖,会触发ModuleNotFoundError导致脚本失败。

解决方案

1. 修复代码中的潜在错误

修改train.py,添加参数校验、异常处理和日志输出:

#!/usr/bin/env python
# coding: utf-8

from azure.ai.ml import MLClient
from azure.ai.ml.entities import Data
from azure.ai.ml.constants import AssetTypes
from azure.identity import DefaultAzureCredential
import sys
import traceback

from azure.storage.blob import BlobServiceClient, BlobClient, ContainerClient, generate_blob_sas, BlobSasPermissions
from azureml.core import Workspace, Dataset
from datetime import datetime, timedelta
from azureml.core import Workspace, Datastore, Dataset
import os

import pandas as pd
import joblib
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error
from sklearn.metrics import accuracy_score

account_name = "mloptestsa"
container_name = "mloptestcontainer"

def get_into_container_blob_storage(connection_string, container_name):
    blob_service_client = BlobServiceClient.from_connection_string(connection_string)
    return blob_service_client.get_container_client(container_name)

def generate_sas_blob_file_with_url(account_name, container_name, blob_name, account_key):
    sas = generate_blob_sas(
        account_name=account_name,
        container_name=container_name,
        blob_name=blob_name,
        account_key=account_key,
        permission=BlobSasPermissions(read=True),
        expiry=datetime.utcnow() + timedelta(hours=1)
    )
    return f'https://{account_name}.blob.core.windows.net/{container_name}/{blob_name}?{sas}'

def get_dataframe_blob_file(connection_string, container_name, account_key):
    container_client = get_into_container_blob_storage(connection_string, container_name)
    blob_target = None
    for blob_i in container_client.list_blobs():
        if blob_i.name == "traindata/car_data.csv":
            blob_target = blob_i.name
            break
    if not blob_target:
        raise FileNotFoundError("未找到traindata/car_data.csv文件")
    
    sas_url = generate_sas_blob_file_with_url(account_name, container_name, blob_target, account_key)
    return pd.read_csv(sas_url, skiprows=1)
    
def filter_data(df): 
    df.columns = ['name', 'year', 'selling_price', 'km_driven', 'fuel', 'seller_type', 'transmission', 'owner']
    df_filtered_columns = df.drop(columns=['fuel','seller_type','owner','transmission'])
    df_filtered_columns['name'], _ = pd.factorize(df_filtered_columns['name'])
    return df_filtered_columns.dropna()

def split_data(df_filtered):
    X = df_filtered.drop(columns=['selling_price']).values
    y = df_filtered['selling_price'].values
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
    return {"train": {"X": X_train, "y": y_train}, "test": {"X": X_test, "y": y_test}}, y_test

def train_model(data):
    model = LinearRegression()
    model.fit(data["train"]["X"], data["train"]["y"])
    return model

def get_model_metrics(model, data, test_y):
    preds = model.predict(data["test"]["X"])
    return {"mse": mean_squared_error(preds, test_y)}

def save_model(model, connection_string, container_name):
    model_name = "price_car_data.pkl"
    joblib.dump(model, filename=model_name)
    container_client = get_into_container_blob_storage(connection_string, container_name)
    blob_client = container_client.get_blob_client(blob=f"model_output/{model_name}")
    with open(model_name, "rb") as data:
        blob_client.upload_blob(data, overwrite=True)
    os.remove(model_name)

def main(connection_string, container_name, account_key):
    df = get_dataframe_blob_file(connection_string, container_name, account_key)
    df_filtered = filter_data(df)
    data, test_y = split_data(df_filtered)
    model = train_model(data)
    metrics = get_model_metrics(model, data, test_y)
    save_model(model, connection_string, container_name)
    print(f"训练完成,指标:{metrics}")

if __name__ == '__main__':
    try:
        if len(sys.argv) < 2:
            raise ValueError("缺少存储账户密钥参数")
        account_key = sys.argv[1]
        connection_string = f"DefaultEndpointsProtocol=https;AccountName={account_name};AccountKey={account_key};EndpointSuffix=core.windows.net"
        main(connection_string, container_name, account_key)
        sys.exit(0)
    except Exception as e:
        print(f"脚本执行失败:{str(e)}")
        traceback.print_exc()
        sys.exit(1)

2. 确保Batch池环境安装依赖

  • 在Batch池的启动任务中添加依赖安装命令:
    pip install pandas scikit-learn azure-storage-blob azure-ai-ml azureml-core
    
  • 或创建requirements.txt文件打包上传,启动任务执行pip install -r requirements.txt

3. 配置ADF任务日志收集

在ADF自定义服务任务中,启用日志存储到Blob容器,通过日志查看脚本输出的详细错误信息,定位具体失败原因。

4. 验证参数传递正确性

手动在本地测试脚本运行,确认账户密钥参数有效:

python train.py <你的存储账户密钥>

内容的提问来源于stack exchange,提问作者DevOps Lover - Hoang An

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 01:20:59