You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Kubernetes上运行Spark 2.3.0使用远程Jar时出错,请求解决方法

Hey there, let’s troubleshoot your failing Spark Pi job on Minikube together— I’ve dealt with this exact scenario a few times, so here’s what to check and fix:

Common Causes & Fixes for Spark Pod Failure on Minikube

1. You’re Using a Placeholder for the Spark K8s Image (Biggest Culprit)

The <default-spark-k8s-image-build> in your command is just a placeholder— Kubernetes can’t pull an image that doesn’t exist. Here’s how to fix this:

  • Option 1: Build your own Spark K8s image
    Use Spark’s built-in tool to create an image tailored for Kubernetes:
    # First, set your Minikube Docker environment so the image is available locally
    eval $(minikube docker-env)
    # Build the image (replace v3.5.0 with your Spark version)
    ./bin/docker-image-tool.sh -r local -t v3.5.0 build
    
    Then update your spark-submit image parameter to local/spark:v3.5.0.
  • Option 2: Use an official pre-built image
    If you don’t want to build locally, use a verified image like apache/spark:3.5.0 (make sure Minikube can access it— run minikube docker pull apache/spark:3.5.0 first).

2. Your Remote JAR URL is Unreachable

The <https://remote-location-with-spark-example-jar> might not be accessible from inside the Minikube cluster. Try these steps:

  • Test if Minikube can reach the URL:
    minikube curl https://remote-location-with-spark-example-jar
    
  • If it fails, switch to using the JAR bundled inside your Spark image (no remote pull needed) by changing the JAR path to local:///opt/spark/examples/jars/spark-examples_*.jar in your spark-submit command.

3. The spark Namespace Doesn’t Exist (Or Permissions Are Missing)

You specified --conf spark.kubernetes.namespace=spark, but if that namespace isn’t created, your Pod will fail to schedule.

  • Create the namespace first:
    kubectl create namespace spark
    
  • Also, Spark needs proper RBAC permissions to run in Kubernetes. Create a service account with edit access:
    apiVersion: v1
    kind: ServiceAccount
    metadata:
      name: spark
      namespace: spark
    ---
    apiVersion: rbac.authorization.k8s.io/v1
    kind: ClusterRoleBinding
    metadata:
      name: spark-role
    subjects:
    - kind: ServiceAccount
      name: spark
      namespace: spark
    roleRef:
      kind: ClusterRole
      name: edit
      apiGroup: rbac.authorization.k8s.io
    
    Apply this with kubectl apply -f spark-rbac.yaml, then add this to your spark-submit command:
    --conf spark.kubernetes.authenticate.driver.serviceAccountName=spark
    

4. Minikube Doesn’t Have Enough Resources

Spark needs a decent amount of memory/CPU to run, and Minikube’s default settings are often too low.

  • Check your current Minikube memory:
    minikube config get memory
    
  • If it’s less than 4GB, increase it and restart Minikube:
    minikube config set memory 4096
    minikube stop && minikube start
    
  • Add resource limits to your Spark job to avoid overloading the cluster:
    --conf spark.driver.memory=1g \
    --conf spark.executor.memory=1g \
    --conf spark.driver.cores=1 \
    --conf spark.executor.cores=1
    

5. Fix Your Command Structure

A quick note: You combined minikube start and spark-submit in one line— this won’t work! minikube start initializes the cluster, so you need to run it separately first, wait for it to finish, then execute your spark-submit command.

Debugging with kubectl describe

You mentioned using kubectl describe— focus on these key sections to pinpoint the issue:

  • Events: Look for errors like FailedPullImage, ErrImagePull, or CrashLoopBackOff— these will tell you exactly why the Pod failed to start.
  • Containers > State: Check the "Waiting" or "Terminated" state details for specific error messages.
  • You can also pull Pod logs (if the container started at all) with:
    kubectl logs <your-pod-name> -n spark
    

Corrected spark-submit Example

Here’s a full, working command that addresses all the above issues:

# 1. Start Minikube (run this first!)
minikube start
# 2. Set Docker environment to use Minikube's local registry
eval $(minikube docker-env)
# 3. Build Spark image (skip if using official pre-built image)
./bin/docker-image-tool.sh -r local -t v3.5.0 build
# 4. Create spark namespace and RBAC (run once)
kubectl create namespace spark
kubectl apply -f spark-rbac.yaml
# 5. Submit the Spark job
./bin/spark-submit \
  --master k8s://https://$(minikube ip):8443 \
  --deploy-mode cluster \
  --name spark-pi \
  --class org.apache.spark.examples.SparkPi \
  --conf spark.executor.instances=1 \
  --conf spark.kubernetes.container.image=local/spark:v3.5.0 \
  --conf spark.kubernetes.namespace=spark \
  --conf spark.kubernetes.authenticate.driver.serviceAccountName=spark \
  --conf spark.driver.memory=1g \
  --conf spark.executor.memory=1g \
  local:///opt/spark/examples/jars/spark-examples_*.jar

内容的提问来源于stack exchange,提问作者Timmeh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:01:33