EKS集群中StatefulSet部署的CockroachDB每日备份至S3的最优方案咨询
Hey folks, let me share the optimal approach I’ve implemented for daily backups of CockroachDB running as a StatefulSet on EKS, with backups pushed to S3. This setup is Kubernetes-native, secure, and proven reliable in production environments.
First, make sure you have these in place before starting:
- A healthy EKS cluster with your CockroachDB StatefulSet deployed (confirm all nodes are reachable and functional)
- An S3 bucket created for backups, with proper IAM permissions configured for write/list access
kubectlconfigured to access your EKS cluster- Matching CockroachDB version (use the same image tag as your running cluster for backup commands)
CockroachDB has a built-in BACKUP SQL command that directly supports writing to S3. Pairing this with a Kubernetes CronJob is the cleanest, most maintainable approach since it leverages native tooling without extra dependencies.
Step 1: Set Up Secure S3 Access with IRSA
For EKS, IAM Roles for Service Accounts (IRSA) is the gold standard for securing S3 access (no hardcoded credentials!). Here's what to do:
- Create an IAM policy that allows
s3:PutObject,s3:ListBucket, ands3:GetBucketLocationpermissions for your target S3 bucket. - Create an IAM role linked to a Kubernetes ServiceAccount (SA) in your CockroachDB namespace.
- Attach the policy to the IAM role, and ensure your CronJob uses this SA.
If IRSA isn't an option for your setup, you can store AWS credentials in a Kubernetes Secret and mount it to the CronJob pod—but IRSA is always preferred for security.
Step 2: Deploy the Backup CronJob
Create a cronjob.yaml file with the following content (replace placeholders with your values):
apiVersion: batch/v1 kind: CronJob metadata: name: cockroachdb-daily-backup namespace: cockroachdb # Replace with your CockroachDB namespace spec: schedule: "0 2 * * *" # Run daily at 2 AM UTC; adjust to your desired timezone concurrencyPolicy: Forbid # Prevent overlapping backup jobs jobTemplate: spec: template: spec: serviceAccountName: cockroachdb-backup-sa # Your IRSA-linked SA containers: - name: cockroachdb-backup image: cockroachdb/cockroach:v23.1.10 # Match your cluster's version command: - /bin/bash - -c - | # Connect to the CockroachDB public service and run backup cockroach sql --host=cockroachdb-public.cockroachdb.svc.cluster.local --port=26257 --user=root <<EOF BACKUP INTO 's3://your-backup-bucket/cockroachdb/daily/{date}' WITH revision_history; EOF env: - name: AWS_REGION value: "us-east-1" # Replace with your S3 bucket's region restartPolicy: OnFailure # Only restart if the backup fails
Key notes:
- The
{date}placeholder in the S3 path is auto-replaced by CockroachDB with the backup timestamp, making it easy to identify daily backups. WITH revision_historyadds point-in-time recovery capability—leave this out only if you don't need historical recovery.concurrencyPolicy: Forbidensures you don't run multiple backups at once if a previous job runs longer than expected.
Step 3: Test the Backup
Manually trigger a test job to verify everything works:
kubectl create job --from=cronjob/cockroachdb-daily-backup test-cockroach-backup -n cockroachdb
Check the job logs to confirm success:
kubectl logs test-cockroach-backup-<pod-suffix> -n cockroachdb
You should see a message like Backup completed successfully—then head to your S3 bucket to confirm the backup files were created.
If you need custom backup logic (like pre/post backup scripts, or incremental backups), you can add a sidecar container to your CockroachDB StatefulSet. However, this is less efficient than a CronJob because you don't need every pod to run backups—CockroachDB backups are cluster-wide, so one instance is enough.
Example sidecar snippet for your StatefulSet:
containers: - name: backup-sidecar image: amazon/aws-cli:latest command: ["bash", "-c"] args: - | while true; do DATE=$(date +%Y-%m-%d) # Connect to the local CockroachDB instance cockroach sql --host=localhost --port=26257 --user=root -e "BACKUP INTO 's3://your-backup-bucket/cockroachdb/daily/$DATE'"; sleep 86400; # Wait 24 hours before next backup done env: - name: AWS_REGION value: "us-east-1" volumeMounts: - name: cockroachdb-certs mountPath: /cockroach/cockroach-certs # Mount TLS certs if your cluster uses encryption
Don't skip these steps:
- S3 Lifecycle Rules: Set up a lifecycle rule in your S3 bucket to automatically delete old backups (e.g., delete backups older than 30 days) to avoid unnecessary storage costs.
- Regular Restore Testing: Periodically test restoring from a backup to a temporary CockroachDB instance. This ensures your backups are valid and you know how to recover data when needed. Example restore command:
cockroach sql --host=temp-cockroach-cluster --port=26257 --user=root <<EOF RESTORE FROM 's3://your-backup-bucket/cockroachdb/daily/2024-05-20' WITH into_db='restored_data'; EOF
内容的提问来源于stack exchange,提问作者shufilkhan

