在Kubernetes上运行Spark 2.3.0使用远程Jar时出错,请求解决方法
Hey there, let’s troubleshoot your failing Spark Pi job on Minikube together— I’ve dealt with this exact scenario a few times, so here’s what to check and fix:
1. You’re Using a Placeholder for the Spark K8s Image (Biggest Culprit)
The <default-spark-k8s-image-build> in your command is just a placeholder— Kubernetes can’t pull an image that doesn’t exist. Here’s how to fix this:
- Option 1: Build your own Spark K8s image
Use Spark’s built-in tool to create an image tailored for Kubernetes:
Then update your# First, set your Minikube Docker environment so the image is available locally eval $(minikube docker-env) # Build the image (replace v3.5.0 with your Spark version) ./bin/docker-image-tool.sh -r local -t v3.5.0 buildspark-submitimage parameter tolocal/spark:v3.5.0. - Option 2: Use an official pre-built image
If you don’t want to build locally, use a verified image likeapache/spark:3.5.0(make sure Minikube can access it— runminikube docker pull apache/spark:3.5.0first).
2. Your Remote JAR URL is Unreachable
The <https://remote-location-with-spark-example-jar> might not be accessible from inside the Minikube cluster. Try these steps:
- Test if Minikube can reach the URL:
minikube curl https://remote-location-with-spark-example-jar - If it fails, switch to using the JAR bundled inside your Spark image (no remote pull needed) by changing the JAR path to
local:///opt/spark/examples/jars/spark-examples_*.jarin yourspark-submitcommand.
3. The spark Namespace Doesn’t Exist (Or Permissions Are Missing)
You specified --conf spark.kubernetes.namespace=spark, but if that namespace isn’t created, your Pod will fail to schedule.
- Create the namespace first:
kubectl create namespace spark - Also, Spark needs proper RBAC permissions to run in Kubernetes. Create a service account with edit access:
Apply this withapiVersion: v1 kind: ServiceAccount metadata: name: spark namespace: spark --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRoleBinding metadata: name: spark-role subjects: - kind: ServiceAccount name: spark namespace: spark roleRef: kind: ClusterRole name: edit apiGroup: rbac.authorization.k8s.iokubectl apply -f spark-rbac.yaml, then add this to yourspark-submitcommand:--conf spark.kubernetes.authenticate.driver.serviceAccountName=spark
4. Minikube Doesn’t Have Enough Resources
Spark needs a decent amount of memory/CPU to run, and Minikube’s default settings are often too low.
- Check your current Minikube memory:
minikube config get memory - If it’s less than 4GB, increase it and restart Minikube:
minikube config set memory 4096 minikube stop && minikube start - Add resource limits to your Spark job to avoid overloading the cluster:
--conf spark.driver.memory=1g \ --conf spark.executor.memory=1g \ --conf spark.driver.cores=1 \ --conf spark.executor.cores=1
5. Fix Your Command Structure
A quick note: You combined minikube start and spark-submit in one line— this won’t work! minikube start initializes the cluster, so you need to run it separately first, wait for it to finish, then execute your spark-submit command.
Debugging with kubectl describe
You mentioned using kubectl describe— focus on these key sections to pinpoint the issue:
- Events: Look for errors like
FailedPullImage,ErrImagePull, orCrashLoopBackOff— these will tell you exactly why the Pod failed to start. - Containers > State: Check the "Waiting" or "Terminated" state details for specific error messages.
- You can also pull Pod logs (if the container started at all) with:
kubectl logs <your-pod-name> -n spark
Corrected spark-submit Example
Here’s a full, working command that addresses all the above issues:
# 1. Start Minikube (run this first!) minikube start # 2. Set Docker environment to use Minikube's local registry eval $(minikube docker-env) # 3. Build Spark image (skip if using official pre-built image) ./bin/docker-image-tool.sh -r local -t v3.5.0 build # 4. Create spark namespace and RBAC (run once) kubectl create namespace spark kubectl apply -f spark-rbac.yaml # 5. Submit the Spark job ./bin/spark-submit \ --master k8s://https://$(minikube ip):8443 \ --deploy-mode cluster \ --name spark-pi \ --class org.apache.spark.examples.SparkPi \ --conf spark.executor.instances=1 \ --conf spark.kubernetes.container.image=local/spark:v3.5.0 \ --conf spark.kubernetes.namespace=spark \ --conf spark.kubernetes.authenticate.driver.serviceAccountName=spark \ --conf spark.driver.memory=1g \ --conf spark.executor.memory=1g \ local:///opt/spark/examples/jars/spark-examples_*.jar
内容的提问来源于stack exchange,提问作者Timmeh

