You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

部署于Google Cloud ML-Engine的Scikit-learn模型预测及欺诈检测模型咨询

Hey Abdul, let's walk through deploying your Scikit-learn fraud detection models (Isolation Forest and Local Outlier Factor) to Google Cloud ML-Engine step by step. I'll cover everything from saving your model to testing the prediction service, with notes specific to your anomaly detection use case.

Deploying Scikit-learn Anomaly Detection Models to Google Cloud ML-Engine

1. Save Your Model Properly

First, you need to save your trained model so ML-Engine can load it later. Scikit-learn recommends using joblib (more efficient than pickle for large models). Here's how to update your code to save the Isolation Forest model:

from sklearn.metrics import classification_report, accuracy_score
from sklearn.ensemble import IsolationForest
from sklearn.neighbors import LocalOutlierFactor
import joblib
import os

# Your existing code
state = 1
# Assume you've defined X and outlier_fraction already
classifiers = {
    "Isolation Forest": IsolationForest(
        max_samples=len(X), 
        contamination=outlier_fraction, 
        random_state=state
    ),
    "Local Outlier Factor": LocalOutlierFactor(
        n_neighbors=20,  # Fill in your missing parameter here
        contamination=outlier_fraction
    )
}

# Train the Isolation Forest (LOOF has special deployment considerations)
model = classifiers["Isolation Forest"]
model.fit(X)

# Create a directory to store the model (make sure it exists)
model_dir = "./model"
os.makedirs(model_dir, exist_ok=True)

# Save the model
model_path = os.path.join(model_dir, "model.joblib")
joblib.dump(model, model_path)

Important Note on LOOF: Local Outlier Factor is a local anomaly detector—it uses the training dataset to compute outlier scores for new data. You can't deploy it standalone like Isolation Forest. If you must use LOOF, you'll need to package the training data with the model and load both in your prediction code. For most production use cases, Isolation Forest is the better choice here since it can predict without relying on the original training data.

2. Prepare Deployment Configuration Files

ML-Engine needs two key files to run your model as a service: setup.py (for dependencies) and predict.py (to handle prediction requests).

setup.py

This file tells ML-Engine which Python packages to install:

from setuptools import setup

REQUIRED_PACKAGES = [
    'scikit-learn>=0.24.0',
    'joblib>=1.0.0'
]

setup(
    name="fraud_detection_model",
    version="0.1",
    install_requires=REQUIRED_PACKAGES,
    packages=["."]
)

predict.py

This is the core file that loads your model and processes incoming prediction requests. It needs a predict function that accepts a request object and returns results:

import joblib
import numpy as np
import os

# Load the model (ML-Engine sets the MODEL_DIR environment variable)
MODEL_DIR = os.environ.get('MODEL_DIR', './model')
model = joblib.load(os.path.join(MODEL_DIR, 'model.joblib'))

def predict(request):
    # Parse JSON request data
    request_json = request.get_json()
    if not request_json or 'data' not in request_json:
        return {"error": "Missing 'data' field in request body"}
    
    # Convert input to a numpy array (matches what the model expects)
    input_data = np.array(request_json['data'])
    
    # Run prediction: Isolation Forest returns -1 for fraud (anomalies) and 1 for normal
    predictions = model.predict(input_data)
    
    # Format results to be human-readable
    results = []
    for pred in predictions:
        results.append({
            "prediction": "fraud" if pred == -1 else "normal",
            "raw_score": pred
        })
    
    return {"predictions": results}

3. Deploy to Google Cloud ML-Engine

Step 1: Upload Files to Google Cloud Storage (GCS)

First, upload your model directory, setup.py, and predict.py to a GCS bucket (create one if you don't have it):

# Replace [YOUR_BUCKET_NAME] with your actual GCS bucket name
gsutil cp -r ./model gs://[YOUR_BUCKET_NAME]/
gsutil cp setup.py gs://[YOUR_BUCKET_NAME]/
gsutil cp predict.py gs://[YOUR_BUCKET_NAME]/

Step 2: Create Model and Version

Use the gcloud CLI to create your model and a deployable version:

# Create the model (replace us-central1 with your preferred region)
gcloud ai-platform models create fraud_detection --region us-central1

# Create a version (match runtime/python versions to your scikit-learn version)
gcloud ai-platform versions create v1 \
    --model fraud_detection \
    --region us-central1 \
    --origin gs://[YOUR_BUCKET_NAME]/ \
    --runtime-version 2.7 \
    --framework scikit-learn \
    --python-version 3.7

Tip: Check Google's docs to confirm runtime/python version compatibility for your scikit-learn version—this avoids deployment errors.

4. Test the Prediction Service

Let's verify your model works as expected.

Option 1: Use gcloud CLI

Create a test JSON file test_request.json with sample input:

{
    "data": [[1.2, 3.4, 5.6, 7.8], [9.0, 8.7, 6.5, 4.3]]
}

Run the prediction command:

gcloud ai-platform predict --model fraud_detection --version v1 --json-request test_request.json

You should get a response like this:

{
  "predictions": [
    {"prediction": "normal", "raw_score": 1},
    {"prediction": "fraud", "raw_score": -1}
  ]
}

Option 2: Use HTTP Request

You can also send a POST request using curl (replace [YOUR_PROJECT_ID] with your GCP project ID):

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  https://us-central1-ml.googleapis.com/v1/projects/[YOUR_PROJECT_ID]/models/fraud_detection/versions/v1:predict \
  -d @test_request.json

5. Key Production Considerations

  • Feature Consistency: Make sure the input features in prediction requests are in the exact same order as your training data—mixing up features will break predictions.
  • LOOF Workaround: If you need to deploy LOOF, save your training data alongside the model and load it in predict.py (use joblib.dump to save the training array). Then use model.fit_predict or calculate scores manually, but this is less efficient than Isolation Forest.
  • Scaling: If you expect high traffic, adjust the machine type when creating the model version (e.g., --machine-type n1-standard-4) to handle more requests.

内容的提问来源于stack exchange,提问作者Abdul Rehman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:05:34