AKS集群中基于Azure Blob Storage创建PV部署Prometheus问题咨询
Hey there! Let's work through getting your Prometheus persistence sorted on AKS. Since you've already got your Azure Blob Storage set up and a Persistent Volume (PV) created, let's dive into the most common issues and fixes step by step.
First up, let's make sure your PV is properly configured for Azure Blob/File storage (Blob is often used via Azure Files for block storage needs like Prometheus). Here's a checklist:
Ensure your PV uses the correct Azure storage driver: For Azure Files (which works well for Prometheus), your PV spec should include the
azureFilesection. Example valid PV YAML:apiVersion: v1 kind: PersistentVolume metadata: name: prometheus-pv labels: app: prometheus # Add a label to help PVC match this PV spec: capacity: storage: 10Gi # Match this to what you set in your Values.yaml accessModes: - ReadWriteOnce # Prometheus only needs single-node write access persistentVolumeReclaimPolicy: Retain azureFile: secretName: azure-storage-secret # Secret with your storage account credentials shareName: prometheus-data # The file share you created in ABC-BLOB-STORAGE readOnly: false mountOptions: - dir_mode=0777 - file_mode=0777 - uid=65534 # Critical: Prometheus runs as the `nobody` user (UID 65534) - gid=65534 # This fixes permission denied errorsDouble-check your storage secret: The secret referenced in the PV must contain your ABC-BLOB-STORAGE account name and key. If you haven't created it yet, run this:
kubectl create secret generic azure-storage-secret \ --from-literal=azurestorageaccountname=ABC-BLOB-STORAGE \ --from-literal=azurestorageaccountkey=<your-storage-account-access-key>Make sure the secret is in the same namespace where you're installing Prometheus (or use a cluster-scoped secret if needed).
Next, your Helm Values.yaml needs to correctly reference the PV (or the PVC that binds to it). Here's how to configure the server.persistence section (Prometheus's main storage component):
Option 1: Use an existing PVC (if you created one to bind your PV)
If you already have a PVC bound to your PV, point Helm to it directly:
server: persistence: enabled: true existingClaim: prometheus-pvc # Name of your pre-created PVC accessModes: - ReadWriteOnce size: 10Gi # Match your PV's capacity
Option 2: Let Helm create a PVC that binds to your static PV
If you want Helm to create the PVC, add a selector to match the labels on your PV:
server: persistence: enabled: true storageClass: "" # Leave empty to avoid dynamic storage provisioning selector: matchLabels: app: prometheus # Match the label you added to your PV accessModes: - ReadWriteOnce size: 10Gi # Must be <= your PV's capacity
Here are the most frequent issues people hit with this setup:
- Permission denied in Prometheus logs: This is almost always due to not setting the
uid/gidin your PV'smountOptions. Prometheus runs asnobody(UID 65534), so the mounted storage needs to be writable by that user. - PV/PVC binding failure: Check that:
- The PV and PVC have matching
accessModes(e.g., bothReadWriteOnce). - The PVC's
sizeis not larger than the PV'scapacity. - The selector labels in the PVC (or Values.yaml) exactly match the labels on your PV.
- The PV and PVC have matching
- Storage account access issues: Ensure your AKS cluster's service principal (or managed identity) has the Storage File Data SMB Contributor role assigned to your ABC-BLOB-STORAGE account. You can set this via the Azure Portal under the storage account's "Access control (IAM)" section.
Once your PV and Values.yaml are sorted, install Prometheus with:
helm install prometheus stable/prometheus \ --values ./values.yaml \ --namespace monitoring \ --create-namespace
If you're upgrading an existing installation, use:
helm upgrade prometheus stable/prometheus \ --values ./values.yaml \ --namespace monitoring
After installation, check that your PV and PVC are bound:
kubectl get pv,pvc -n monitoring
You should see Bound in the STATUS column for both. Then check the Prometheus pod status:
kubectl get pods -n monitoring
If the pod is stuck in Pending or CrashLoopBackOff, pull the logs to diagnose further:
kubectl logs <prometheus-server-pod-name> -n monitoring
内容的提问来源于stack exchange,提问作者frictionlesspulley

