Azure ML实验服务部署添额外脚本及JSON数据模型训练技术咨询
Hey there, let's break down your two Azure ML questions with practical, actionable steps based on my hands-on experience:
The key to packaging extra scripts (like preprocessing utilities or helper functions) is to ensure they're included in the deployment bundle. Here's how to do it properly:
Use the
source_directoryParameter in Deployment
When setting up your inference configuration, specify the folder that contains all your required scripts (including your mainscore.pyinference script). Azure ML will package every file in this directory and make them available to your deployed service.Example code snippet:
from azureml.core import Model, Environment, InferenceConfig from azureml.core.webservice import AciWebservice # Load your trained model from the workspace model = Model(workspace, name="your-trained-model") # Define your environment (use existing or custom conda config) env = Environment.from_conda_specification(name="model-env", file_path="./conda.yml") # Set up inference config with your script directory inference_config = InferenceConfig( environment=env, source_directory="./model-scripts", # All files here get packaged entry_script="score.py" # Your main inference entry point ) # Deploy the service deployment_config = AciWebservice.deploy_configuration(cpu_cores=1, memory_gb=1) service = Model.deploy( workspace=workspace, name="your-deployed-service", models=[model], inference_config=inference_config, deployment_config=deployment_config ) service.wait_for_deployment(show_output=True)Once deployed, you can import your extra scripts directly in
score.py(e.g.,from data_transform import preprocess_raw_data).Handle Dependencies for Custom Scripts
If your extra scripts rely on specific Python packages, make sure to list them in yourconda.ymlfile. For example:name: model-env channels: - conda-forge dependencies: - python=3.8 - scikit-learn=1.0.2 - pip: - azureml-defaults - pandas=1.4.2
Since you already have the JSON-to-feature conversion code, the goal is to integrate this logic into your inference pipeline and ensure end-to-end consistency. Here's how to make it work:
Embed Preprocessing Logic in Your Inference Flow
Move your existing data transformation function into a reusable script (e.g.,data_utils.py) and call it from yourscore.pyduring inference. This ensures the same preprocessing logic is used for both training and prediction.Example
data_utils.py:def transform_json_to_features(raw_json_data): # Your existing feature extraction logic here # Example: extract fields, compute derived features, format for model input processed_features = [] for record in raw_json_data: feature_set = [ record["user_age"], record["transaction_amount"] * 0.75, record["transaction_frequency"] ] processed_features.append(feature_set) return processed_featuresExample
score.py:import json import joblib from data_utils import transform_json_to_features def init(): global model # Load your trained model from the deployment bundle model = joblib.load("trained_model.pkl") def run(raw_data): try: # Parse incoming JSON input input_data = json.loads(raw_data) # Apply the same preprocessing as training model_input = transform_json_to_features(input_data) # Run prediction predictions = model.predict(model_input) # Format output for the client return {"predictions": predictions.tolist()} except Exception as e: return {"error": str(e)}Enforce Input Schema Consistency
To ensure incoming requests match your expected JSON structure, define a schema file (schema.json) and attach it to your inference configuration. Azure ML will automatically validate inputs against this schema.Example
schema.json:{ "$schema": "http://json-schema.org/draft-04/schema#", "type": "array", "items": { "type": "object", "properties": { "user_age": {"type": "integer"}, "transaction_amount": {"type": "number"}, "transaction_frequency": {"type": "integer"} }, "required": ["user_age", "transaction_amount", "transaction_frequency"] } }Update your inference config to include the schema:
inference_config = InferenceConfig( environment=env, source_directory="./model-scripts", entry_script="score.py", schema_file="./schema.json" )Test Locally Before Cloud Deployment
Debugging locally saves time—use Azure ML's local deployment option to validate your pipeline:from azureml.core.webservice import LocalWebservice local_config = LocalWebservice.deploy_configuration(port=8000) local_service = Model.deploy( workspace=workspace, name="local-test-service", models=[model], inference_config=inference_config, deployment_config=local_config ) local_service.wait_for_deployment() # Test with sample JSON input sample_input = json.dumps([ {"user_age": 30, "transaction_amount": 150.50, "transaction_frequency": 5}, {"user_age": 45, "transaction_amount": 200.00, "transaction_frequency": 3} ]) response = local_service.run(sample_input) print(response)
内容的提问来源于stack exchange,提问作者Emil L

