You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure ML Studio中运行Spark作业可行性咨询及问题求助

Azure ML Studio: Running Spark Jobs & Deploying Web Services

Hey there! Let's tackle your questions about running Spark scripts in Azure ML Studio and deploying your model as a web service:

First, the core issue: Can you run Spark jobs directly in Azure ML Studio (classic)?

Short answer: No, the classic Azure ML Studio's "Execute Python Script" module doesn't natively support a Spark runtime environment—that's exactly why you're hitting import errors for Spark libraries. The module runs in a standard Python environment without Spark's distributed framework or pre-installed dependencies.

So how do you run Spark jobs with Azure ML?

You have two solid paths here:

  • Use the modern Azure Machine Learning Workspace: This is the recommended approach. It fully supports Spark compute clusters/instances. You can run your Spark scripts directly in notebooks, set up Spark-based training jobs, and natively import pyspark and related libraries without hassle.
  • Integrate with Azure Databricks (if you need to stick with classic ML Studio):
    • Run your Spark workloads in Azure Databricks (Azure's managed Spark service), save outputs to Azure Blob Storage.
    • Then pull those processed results into classic ML Studio for further modeling, or use the "Execute Python Script" module to call Databricks APIs and trigger Spark jobs indirectly.

If you just need to test small-scale Spark logic in classic ML Studio, you could manually upload Spark-related wheel packages to your ML Studio dataset, then add the package path to your script with:

import sys
sys.path.append(".\\Script Bundle\\your-spark-package")
import pyspark

Note: This only lets you run local, non-distributed Spark code—it won't work for large-scale distributed jobs, and you'll likely hit compatibility issues.

Deploying your Spark model as a web service

  • Modern Azure ML Workspace: This is the way to go. Save your Spark MLlib model in a compatible format (like MLflow), then deploy it directly as a real-time web service or batch inference service. The workspace handles all the runtime setup for Spark models.
  • Classic ML Studio: You can't deploy Spark models directly here. You'd need to convert your Spark inference logic to use standard Python (e.g., replace Spark DataFrame operations with Pandas) and wrap that in the "Execute Python Script" module, then deploy the pipeline as a web service. This only works for small-scale inference tasks.

内容的提问来源于stack exchange,提问作者Ravi Kiran G

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:44:47