You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

咨询Spark 2.3在Kubernetes 1.8.*集群的REST提交接口及AWS Lambda调用方案

回答

Absolutely! Kubernetes' API Server provides a full REST interface that you can use to submit Spark 2.3 jobs to your 1.8.* cluster from AWS Lambda—no need to run spark-submit directly in your Lambda function. Here's a breakdown of how to make this work:

1. Understand the Kubernetes API Foundation

Every operation you perform with kubectl or spark-submit (when targeting K8s) ultimately calls the K8s API Server's REST endpoints. For your use case, you'll send authenticated HTTP requests from Lambda to these endpoints to orchestrate your Spark jobs.

2. Two Practical Approaches to Submit Spark Jobs via the API

Option 1: Create a Spark Driver Pod Directly

In Spark 2.3's cluster mode, spark-submit creates a Driver Pod which then spawns Executor Pods. You can replicate this by constructing a Pod specification and sending it to the K8s API:

  • Target the endpoint: POST /api/v1/namespaces/<your-namespace>/pods
  • Your Pod spec needs to include:
    • A Spark 2.3 image compatible with K8s (e.g., spark:2.3.0-kubernetes or a custom build with AWS SDKs for S3 access)
    • The spark-submit command as the container's entrypoint, with --master k8s://<k8s-api-server-url> and your job arguments
    • Necessary environment variables (like AWS_REGION or credentials if using S3)
    • Volume mounts (if your job needs access to configs or data)
    • A Service to enable communication between Driver and Executors

Here's a simplified JSON snippet for the Pod spec (ready for API submission):

{
  "apiVersion": "v1",
  "kind": "Pod",
  "metadata": {
    "name": "spark-driver-myjob",
    "labels": {
      "spark-role": "driver"
    }
  },
  "spec": {
    "containers": [
      {
        "name": "spark-driver",
        "image": "spark:2.3.0-kubernetes",
        "command": [
          "spark-submit",
          "--master", "k8s://https://<k8s-api-server>",
          "--deploy-mode", "cluster",
          "--class", "com.yourcompany.YourSparkJob",
          "s3://your-bucket/jars/your-spark-job.jar"
        ],
        "env": [
          {"name": "AWS_REGION", "value": "us-east-1"}
        ]
      }
    ]
  }
}

Option 2: Use a Kubernetes Job to Run spark-submit

If you prefer to leverage the existing logic of spark-submit, you can create a K8s Job that runs the command for you. This is great if you already have a working spark-submit command and want to reuse it:

  • Target the endpoint: POST /apis/batch/v1/namespaces/<your-namespace>/jobs
  • The Job spec defines a Pod that executes your spark-submit command, and K8s handles managing the Pod's lifecycle (including cleaning up after the job finishes, if configured)

Example Job spec snippet:

{
  "apiVersion": "batch/v1",
  "kind": "Job",
  "metadata": {
    "name": "spark-job-submitter"
  },
  "spec": {
    "template": {
      "spec": {
        "containers": [
          {
            "name": "spark-submit",
            "image": "spark:2.3.0-kubernetes",
            "command": [
              "spark-submit",
              "--master", "k8s://https://<k8s-api-server>",
              "--class", "com.yourcompany.YourSparkJob",
              "s3://your-bucket/jars/your-spark-job.jar"
            ]
          }
        ],
        "restartPolicy": "OnFailure"
      }
    }
  }
}

3. Critical Permission & Network Setup

Since you're on AWS (likely using EKS), here's what you need to configure:

  • Lambda to K8s API Access: Use IAM Roles for Service Accounts (IRSA) to map your Lambda's IAM role to a K8s ServiceAccount with permissions to create Pods/Jobs. This avoids hardcoding credentials.
  • Network Access: Ensure your Lambda is in a VPC that can reach the K8s API Server (either via private endpoint or public endpoint with proper security group rules).
  • K8s RBAC: Grant the mapped ServiceAccount the necessary RBAC roles (e.g., edit role in your target namespace) to create and manage Spark-related resources.

4. Spark 2.3 Specific Notes

Spark 2.3's K8s support was in beta, so keep these in mind:

  • Use a K8s-compatible Spark image (official or custom) that includes the spark-submit binary and any dependencies your job needs.
  • Verify that your K8s cluster's network policy allows communication between the Spark Driver and Executors.
  • For S3 access, ensure the Spark image has the AWS SDKs or use IAM roles for the Spark Pods (via IRSA as well).

内容的提问来源于stack exchange,提问作者shiv455

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:32:12