You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Azure Databricks中使用自定义Docker镜像创建临时作业集群?

Can I use a custom Docker image to create a temporary job cluster in Azure Databricks?

Yes, you can absolutely use a custom Docker image to spin up temporary job clusters in Azure Databricks through the Jobs API. Here's a breakdown of what you need to know, plus targeted feedback on your job definition JSON:

Key Requirements & Notes

  • Spark Version Compatibility: Your custom Docker image must be built to match the spark_version specified in your cluster configuration (in your case, 7.3.x-scala2.12). Databricks provides base images you can extend to ensure compatibility—make sure your image includes all necessary Spark dependencies and Databricks runtime components to avoid startup failures.
  • ACR Authentication: The basic_auth block for your Azure Container Registry (ACR) is correctly structured, but double-check a few details:
    • The <ACR_HOST_NAME> needs the full image path (e.g., myacr.azurecr.io/my-custom-image:latest)
    • The <ACR_USER> should be your ACR registry name
    • The <ACR_TOKEN> must be a valid access token with pull permissions for the target image
  • Instance Pool Considerations: If you’re using an instance pool (instance_pool_id), confirm the pool is configured to allow custom Docker images—some pool settings may restrict external image usage. Alternatively, you can omit the pool and define instance types directly in new_cluster if pool benefits aren’t necessary for your use case.
  • Temporary Cluster Behavior: Since you’re defining new_cluster in your job settings, this will automatically create a temporary cluster that terminates after the job completes—exactly what you’re looking for.

Feedback on Your Job Definition JSON

Your configuration is mostly correct, but there are a few tweaks to avoid validation errors and ensure functionality:

  1. Remove databricks_pool_name: This field isn’t part of the official /jobs/create API schema—you only need instance_pool_id inside new_cluster to link to your pool.
  2. Full Image URL: Update the url in docker_image to include the full image reference (registry host, repository name, and tag) to ensure Databricks can pull the correct image.
  3. Validate Compatibility: Double-check that your custom image is built specifically for 7.3.x-scala2.12—mismatched versions will cause cluster startup failures.

Here’s the adjusted JSON with these fixes:

{ 
  "job_settings": { 
    "name": "job-test", 
    "new_cluster": { 
      "num_workers": 1, 
      "spark_version": "7.3.x-scala2.12", 
      "instance_pool_id": "<INSTANCE_POOL_PLACEHOLDER>", 
      "docker_image": { 
        "url": "<ACR_HOST_NAME>/my-custom-image:tag", 
        "basic_auth": { 
          "username": "<ACR_USER>", 
          "password": "<ACR_TOKEN>" 
        } 
      } 
    }, 
    "max_concurrent_runs": 1, 
    "max_retries": 0, 
    "schedule": { 
      "quartz_cron_expression": "0 0 0 2 * ?", 
      "timezone_id": "UTC" 
    }, 
    "spark_python_task": { 
      "python_file": "dbfs:/poc.py" 
    }, 
    "timeout_seconds": 5400 
  } 
}

Additional Tips

  • Test Image Locally: Before deploying, run basic Spark commands in your image to catch compatibility issues early.
  • Check Cluster Logs: If the cluster fails to start, review the Databricks cluster logs (under "Cluster" > "Logs") for specific errors related to image pulling or runtime initialization.
  • Secure Authentication: For production, consider using Azure service principals instead of basic auth for ACR access—it’s more secure and scalable.

内容的提问来源于stack exchange,提问作者Lorenz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 10:32:46