You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在AWS EKS集群部署Ollama启动失败,求排查及配置修正

问题排查与EKS Ollama GPU服务配置修正

问题根源

  1. Init容器命令逻辑错误:ollama run需要依赖本地运行的Ollama服务,但Init容器中直接执行该命令时,Ollama进程未启动,导致连接失败报错。
  2. GPU调度缺失:改用节点选择器后,未在容器中声明GPU资源限制,无法确保Pod能调用GPU资源,同时之前的调度失败大概率是因为集群未配置NVIDIA设备插件。
  3. 存活探针端口错误:原配置监听80端口,但Ollama默认服务端口为11434,探针会误判服务状态。
  4. HostPath PV适配问题:EKS集群中使用hostPath类型PV会导致数据绑定到单个节点,不符合Kubernetes分布式调度特性,且在托管节点组中存在权限和生命周期风险。

修正后的完整inference.yaml配置

# 动态创建EBS类型PV(适配EKS集群,替代HostPath)
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: ollama-charlie-pvc
spec:
  accessModes:
    - ReadWriteOnce
  resources:
    requests:
      storage: 100Gi
  storageClassName: gp3  # 使用AWS默认gp3存储类,可根据需求修改

---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: ollama-charlie
spec:
  replicas: 1
  selector:
    matchLabels:
      app: ollama-charlie
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  template:
    metadata:
      labels:
        app: ollama-charlie
    spec:
      nodeSelector:
        ollama-charlie-key: ollama-charlie-value
      initContainers:
      - name: download-llm
        image: ollama/ollama
        # 使用ollama pull直接下载模型,无需依赖Ollama服务运行
        command: ["ollama", "pull", "kristada673/solar-10.7b-instruct-v1.0-uncensored"]
        volumeMounts:
        - name: data
          mountPath: /root/.ollama
        # 声明GPU资源,确保Init容器能利用GPU加速模型下载
        resources:
          limits:
            nvidia.com/gpu: 1
      containers:
      - name: ollama-charlie
        image: ollama/ollama
        volumeMounts:
        - name: data
          mountPath: /root/.ollama
        # 声明GPU资源,确保主容器能调用GPU推理
        resources:
          limits:
            nvidia.com/gpu: 1
        livenessProbe:
          tcpSocket:
            port: 11434  # 修正为Ollama默认服务端口
          initialDelaySeconds: 120
          periodSeconds: 30
          failureThreshold: 2
        readinessProbe:
          httpGet:
            path: /api/tags
            port: 11434
          initialDelaySeconds: 60
          periodSeconds: 10
      volumes:
      - name: data
        persistentVolumeClaim:
          claimName: ollama-charlie-pvc
      restartPolicy: Always

---
apiVersion: v1
kind: Service
metadata:
  name: ollama-charlie-service
spec:
  selector:
    app: ollama-charlie
  ports:
    - protocol: TCP
      port: 11434
      targetPort: 11434

核心修改说明

  • 替换Init容器命令:用ollama pull替代ollama run,该命令无需启动Ollama服务即可直接将模型下载到挂载的存储卷,解决连接失败问题。
  • 添加GPU资源声明:在Init容器和主容器中都配置nvidia.com/gpu:1,确保Pod能调度到GPU节点并利用GPU资源。
  • 修复探针配置:将存活探针端口改为11434,新增基于/api/tags接口的就绪探针,更准确地检测服务状态。
  • 替换PV类型:移除HostPath PV,改用EBS动态PV,适配EKS集群的分布式调度特性。

额外检查项

  1. 验证NVIDIA设备插件:执行kubectl get pods -n kube-system,确认nvidia-device-plugin-daemonset类Pod正常运行,否则集群无法识别GPU资源。
  2. 节点标签校验:执行kubectl get nodes --show-labels,确认GPU节点组已正确设置ollama-charlie-key: ollama-charlie-value标签。
  3. 调整超时时间:大模型下载耗时较长,若Init容器超时,可添加terminationGracePeriodSeconds: 1800延长超时时间。

内容的提问来源于stack exchange,提问作者Kristada673

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 09:22:06