咨询Spark 2.3在Kubernetes 1.8.*集群的REST提交接口及AWS Lambda调用方案
Absolutely! Kubernetes' API Server provides a full REST interface that you can use to submit Spark 2.3 jobs to your 1.8.* cluster from AWS Lambda—no need to run spark-submit directly in your Lambda function. Here's a breakdown of how to make this work:
1. Understand the Kubernetes API Foundation
Every operation you perform with kubectl or spark-submit (when targeting K8s) ultimately calls the K8s API Server's REST endpoints. For your use case, you'll send authenticated HTTP requests from Lambda to these endpoints to orchestrate your Spark jobs.
2. Two Practical Approaches to Submit Spark Jobs via the API
Option 1: Create a Spark Driver Pod Directly
In Spark 2.3's cluster mode, spark-submit creates a Driver Pod which then spawns Executor Pods. You can replicate this by constructing a Pod specification and sending it to the K8s API:
- Target the endpoint:
POST /api/v1/namespaces/<your-namespace>/pods - Your Pod spec needs to include:
- A Spark 2.3 image compatible with K8s (e.g.,
spark:2.3.0-kubernetesor a custom build with AWS SDKs for S3 access) - The
spark-submitcommand as the container's entrypoint, with--master k8s://<k8s-api-server-url>and your job arguments - Necessary environment variables (like
AWS_REGIONor credentials if using S3) - Volume mounts (if your job needs access to configs or data)
- A Service to enable communication between Driver and Executors
- A Spark 2.3 image compatible with K8s (e.g.,
Here's a simplified JSON snippet for the Pod spec (ready for API submission):
{ "apiVersion": "v1", "kind": "Pod", "metadata": { "name": "spark-driver-myjob", "labels": { "spark-role": "driver" } }, "spec": { "containers": [ { "name": "spark-driver", "image": "spark:2.3.0-kubernetes", "command": [ "spark-submit", "--master", "k8s://https://<k8s-api-server>", "--deploy-mode", "cluster", "--class", "com.yourcompany.YourSparkJob", "s3://your-bucket/jars/your-spark-job.jar" ], "env": [ {"name": "AWS_REGION", "value": "us-east-1"} ] } ] } }
Option 2: Use a Kubernetes Job to Run spark-submit
If you prefer to leverage the existing logic of spark-submit, you can create a K8s Job that runs the command for you. This is great if you already have a working spark-submit command and want to reuse it:
- Target the endpoint:
POST /apis/batch/v1/namespaces/<your-namespace>/jobs - The Job spec defines a Pod that executes your
spark-submitcommand, and K8s handles managing the Pod's lifecycle (including cleaning up after the job finishes, if configured)
Example Job spec snippet:
{ "apiVersion": "batch/v1", "kind": "Job", "metadata": { "name": "spark-job-submitter" }, "spec": { "template": { "spec": { "containers": [ { "name": "spark-submit", "image": "spark:2.3.0-kubernetes", "command": [ "spark-submit", "--master", "k8s://https://<k8s-api-server>", "--class", "com.yourcompany.YourSparkJob", "s3://your-bucket/jars/your-spark-job.jar" ] } ], "restartPolicy": "OnFailure" } } } }
3. Critical Permission & Network Setup
Since you're on AWS (likely using EKS), here's what you need to configure:
- Lambda to K8s API Access: Use IAM Roles for Service Accounts (IRSA) to map your Lambda's IAM role to a K8s ServiceAccount with permissions to create Pods/Jobs. This avoids hardcoding credentials.
- Network Access: Ensure your Lambda is in a VPC that can reach the K8s API Server (either via private endpoint or public endpoint with proper security group rules).
- K8s RBAC: Grant the mapped ServiceAccount the necessary RBAC roles (e.g.,
editrole in your target namespace) to create and manage Spark-related resources.
4. Spark 2.3 Specific Notes
Spark 2.3's K8s support was in beta, so keep these in mind:
- Use a K8s-compatible Spark image (official or custom) that includes the
spark-submitbinary and any dependencies your job needs. - Verify that your K8s cluster's network policy allows communication between the Spark Driver and Executors.
- For S3 access, ensure the Spark image has the AWS SDKs or use IAM roles for the Spark Pods (via IRSA as well).
内容的提问来源于stack exchange,提问作者shiv455

