部署于Google Cloud ML-Engine的Scikit-learn模型预测及欺诈检测模型咨询
Hey Abdul, let's walk through deploying your Scikit-learn fraud detection models (Isolation Forest and Local Outlier Factor) to Google Cloud ML-Engine step by step. I'll cover everything from saving your model to testing the prediction service, with notes specific to your anomaly detection use case.
1. Save Your Model Properly
First, you need to save your trained model so ML-Engine can load it later. Scikit-learn recommends using joblib (more efficient than pickle for large models). Here's how to update your code to save the Isolation Forest model:
from sklearn.metrics import classification_report, accuracy_score from sklearn.ensemble import IsolationForest from sklearn.neighbors import LocalOutlierFactor import joblib import os # Your existing code state = 1 # Assume you've defined X and outlier_fraction already classifiers = { "Isolation Forest": IsolationForest( max_samples=len(X), contamination=outlier_fraction, random_state=state ), "Local Outlier Factor": LocalOutlierFactor( n_neighbors=20, # Fill in your missing parameter here contamination=outlier_fraction ) } # Train the Isolation Forest (LOOF has special deployment considerations) model = classifiers["Isolation Forest"] model.fit(X) # Create a directory to store the model (make sure it exists) model_dir = "./model" os.makedirs(model_dir, exist_ok=True) # Save the model model_path = os.path.join(model_dir, "model.joblib") joblib.dump(model, model_path)
Important Note on LOOF: Local Outlier Factor is a local anomaly detector—it uses the training dataset to compute outlier scores for new data. You can't deploy it standalone like Isolation Forest. If you must use LOOF, you'll need to package the training data with the model and load both in your prediction code. For most production use cases, Isolation Forest is the better choice here since it can predict without relying on the original training data.
2. Prepare Deployment Configuration Files
ML-Engine needs two key files to run your model as a service: setup.py (for dependencies) and predict.py (to handle prediction requests).
setup.py
This file tells ML-Engine which Python packages to install:
from setuptools import setup REQUIRED_PACKAGES = [ 'scikit-learn>=0.24.0', 'joblib>=1.0.0' ] setup( name="fraud_detection_model", version="0.1", install_requires=REQUIRED_PACKAGES, packages=["."] )
predict.py
This is the core file that loads your model and processes incoming prediction requests. It needs a predict function that accepts a request object and returns results:
import joblib import numpy as np import os # Load the model (ML-Engine sets the MODEL_DIR environment variable) MODEL_DIR = os.environ.get('MODEL_DIR', './model') model = joblib.load(os.path.join(MODEL_DIR, 'model.joblib')) def predict(request): # Parse JSON request data request_json = request.get_json() if not request_json or 'data' not in request_json: return {"error": "Missing 'data' field in request body"} # Convert input to a numpy array (matches what the model expects) input_data = np.array(request_json['data']) # Run prediction: Isolation Forest returns -1 for fraud (anomalies) and 1 for normal predictions = model.predict(input_data) # Format results to be human-readable results = [] for pred in predictions: results.append({ "prediction": "fraud" if pred == -1 else "normal", "raw_score": pred }) return {"predictions": results}
3. Deploy to Google Cloud ML-Engine
Step 1: Upload Files to Google Cloud Storage (GCS)
First, upload your model directory, setup.py, and predict.py to a GCS bucket (create one if you don't have it):
# Replace [YOUR_BUCKET_NAME] with your actual GCS bucket name gsutil cp -r ./model gs://[YOUR_BUCKET_NAME]/ gsutil cp setup.py gs://[YOUR_BUCKET_NAME]/ gsutil cp predict.py gs://[YOUR_BUCKET_NAME]/
Step 2: Create Model and Version
Use the gcloud CLI to create your model and a deployable version:
# Create the model (replace us-central1 with your preferred region) gcloud ai-platform models create fraud_detection --region us-central1 # Create a version (match runtime/python versions to your scikit-learn version) gcloud ai-platform versions create v1 \ --model fraud_detection \ --region us-central1 \ --origin gs://[YOUR_BUCKET_NAME]/ \ --runtime-version 2.7 \ --framework scikit-learn \ --python-version 3.7
Tip: Check Google's docs to confirm runtime/python version compatibility for your scikit-learn version—this avoids deployment errors.
4. Test the Prediction Service
Let's verify your model works as expected.
Option 1: Use gcloud CLI
Create a test JSON file test_request.json with sample input:
{ "data": [[1.2, 3.4, 5.6, 7.8], [9.0, 8.7, 6.5, 4.3]] }
Run the prediction command:
gcloud ai-platform predict --model fraud_detection --version v1 --json-request test_request.json
You should get a response like this:
{ "predictions": [ {"prediction": "normal", "raw_score": 1}, {"prediction": "fraud", "raw_score": -1} ] }
Option 2: Use HTTP Request
You can also send a POST request using curl (replace [YOUR_PROJECT_ID] with your GCP project ID):
curl -X POST \ -H "Authorization: Bearer $(gcloud auth print-access-token)" \ -H "Content-Type: application/json" \ https://us-central1-ml.googleapis.com/v1/projects/[YOUR_PROJECT_ID]/models/fraud_detection/versions/v1:predict \ -d @test_request.json
5. Key Production Considerations
- Feature Consistency: Make sure the input features in prediction requests are in the exact same order as your training data—mixing up features will break predictions.
- LOOF Workaround: If you need to deploy LOOF, save your training data alongside the model and load it in
predict.py(usejoblib.dumpto save the training array). Then usemodel.fit_predictor calculate scores manually, but this is less efficient than Isolation Forest. - Scaling: If you expect high traffic, adjust the machine type when creating the model version (e.g.,
--machine-type n1-standard-4) to handle more requests.
内容的提问来源于stack exchange,提问作者Abdul Rehman

