如何在GKE的COS节点上配置RAID?合并双本地SSD为单目录
Got it, let's walk through exactly how to combine your two local SSDs into a single directory on GKE's Container-Optimized OS (COS) nodes. Since you want this configured at the node level (not per pod), we'll use a node startup script to automate the RAID setup every time a node spins up—perfect for scaling or node replacements.
First, quick context: On COS nodes, each local SSD shows up as an NVMe block device (usually /dev/nvme0n1 and /dev/nvme1n1 for two SSDs), with a default temporary mount at /mnt/disks/ssd0 and /mnt/disks/ssd1. We'll build a RAID 0 array (striped, for maximum performance and combined capacity) using these raw devices.
Step 1: Create the Startup Script
This script handles installing RAID tools, unmounting the default SSD mounts, creating the array, formatting it, and setting up persistent mounting (critical since COS has a read-only root filesystem).
#!/bin/bash # Install mdadm (the tool we need for RAID setup) apt-get update && apt-get install -y mdadm # Unmount the default COS mounts for local SSDs (we need the raw devices for RAID) umount /mnt/disks/ssd0 umount /mnt/disks/ssd1 # Create a RAID 0 array with the two local SSDs mdadm --create /dev/md0 --level=0 --raid-devices=2 /dev/nvme0n1 /dev/nvme1n1 # Format the RAID array with ext4 (a reliable, widely supported filesystem) mkfs.ext4 /dev/md0 # Create a directory to mount our combined RAID volume mkdir -p /mnt/raid-ssd # Mount the RAID array to the new directory mount /dev/md0 /mnt/raid-ssd # Add an entry to /etc/fstab so the mount survives node reboots echo "/dev/md0 /mnt/raid-ssd ext4 defaults 0 0" >> /etc/fstab # Set permissions so pods can access the directory (adjust this to match your workload's user ID if needed) chmod 777 /mnt/raid-ssd
Step 2: Apply the Script to Your Node Pool
You can add this script when creating a new node pool or updating an existing one.
For a New Node Pool
Use the gcloud command to create a node pool with the startup script embedded:
gcloud container node-pools create raid-enabled-pool \ --cluster=YOUR_CLUSTER_NAME \ --zone=YOUR_CLUSTER_ZONE \ --num-nodes=3 \ --machine-type=n2-standard-4 \ --local-ssd-count=2 \ --metadata startup-script='#!/bin/bash apt-get update && apt-get install -y mdadm umount /mnt/disks/ssd0 umount /mnt/disks/ssd1 mdadm --create /dev/md0 --level=0 --raid-devices=2 /dev/nvme0n1 /dev/nvme1n1 mkfs.ext4 /dev/md0 mkdir -p /mnt/raid-ssd mount /dev/md0 /mnt/raid-ssd echo "/dev/md0 /mnt/raid-ssd ext4 defaults 0 0" >> /etc/fstab chmod 777 /mnt/raid-ssd'
For an Existing Node Pool
Update your node pool to include the startup script, then replace existing nodes (since startup scripts only run on node creation):
# Update the node pool with the startup script gcloud container node-pools update YOUR_EXISTING_POOL_NAME \ --cluster=YOUR_CLUSTER_NAME \ --zone=YOUR_CLUSTER_ZONE \ --metadata startup-script='#!/bin/bash apt-get update && apt-get install -y mdadm umount /mnt/disks/ssd0 umount /mnt/disks/ssd1 mdadm --create /dev/md0 --level=0 --raid-devices=2 /dev/nvme0n1 /dev/nvme1n1 mkfs.ext4 /dev/md0 mkdir -p /mnt/raid-ssd mount /dev/md0 /mnt/raid-ssd echo "/dev/md0 /mnt/raid-ssd ext4 defaults 0 0" >> /etc/fstab chmod 777 /mnt/raid-ssd' # Replace existing nodes to apply the new script (scale down then up, or use node replacement commands) gcloud container clusters resize YOUR_CLUSTER_NAME \ --zone=YOUR_CLUSTER_ZONE \ --node-pool=YOUR_EXISTING_POOL_NAME \ --num-nodes=0 gcloud container clusters resize YOUR_CLUSTER_NAME \ --zone=YOUR_CLUSTER_ZONE \ --node-pool=YOUR_EXISTING_POOL_NAME \ --num-nodes=3
Step 3: Access the RAID Volume from Your Pods
To let your workload use the combined SSD directory, mount it as a hostPath volume in your pod spec:
apiVersion: v1 kind: Pod metadata: name: raid-test-pod spec: containers: - name: test-container image: nginx volumeMounts: - name: raid-storage mountPath: /app/data # Path where your workload will access the storage volumes: - name: raid-storage hostPath: path: /mnt/raid-ssd # The RAID directory we created on the node type: Directory
Key Notes
- RAID Level: We used RAID 0 here for maximum performance and capacity, but if you need redundancy (at the cost of halving usable space), swap
--level=0to--level=1in the script for RAID 1 (mirroring). - Ephemeral Storage: Remember that local SSDs are ephemeral—data will be lost when a node is deleted or replaced. If you need persistent storage, use GCP Persistent Disks instead.
- Permissions: The
chmod 777is a quick test setup. For production, adjust permissions to match the user ID your pod runs as (e.g.,chown 1000:1000 /mnt/raid-ssdif your pod uses UID 1000).
内容的提问来源于stack exchange,提问作者Andy

