如何通过Terraform限制GCP Kubernetes集群的磁盘用量
问题描述
背景
我正在使用GCP免费套餐,尝试学习将GitOps与IaC结合应用。计划以Google为云服务商,用Terraform创建Kubernetes集群基础设施,已配置Github Actions实现代码推送时自动应用变更,但遇到如下报错:
│ Error: googleapi: Error 403: Insufficient regional quota to satisfy request: resource "SSD_TOTAL_GB": request requires '300.0' and is short '50.0'. project has a quota of '250.0' with '250.0' available. View and manage quotas at https://console.cloud.google.com/iam-admin/quotas?usage=USED&project=swift-casing-370717., forbidden │ │ with google_container_cluster.primary, │ on main.tf line 26, in resource "google_container_cluster" "primary": │ 26: resource "google_container_cluster" "primary" ***
现有配置
对应的Terraform配置文件:
# https://registry.terraform.io/providers/hashicorp/google/latest/docs provider "google" { project = "redacted" region = "europe-west9" } # https://www.terraform.io/language/settings/backends/gcs terraform { backend "gcs" { bucket = "redacted" prefix = "terraform/state" } required_providers { google = { source = "hashicorp/google" version = "~> 4.0" } } } resource "google_service_account" "default" { account_id = "service-account-id" display_name = "We still use master" } resource "google_container_cluster" "primary" { name = "k8s-cluster" location = "europe-west9" # We can't create a cluster with no node pool defined, but we want to only use # separately managed node pools. So we create the smallest possible default # node pool and immediately delete it. remove_default_node_pool = true initial_node_count = 1 } resource "google_container_node_pool" "primary_preemptible_nodes" { name = "k8s-node-pool" location = "europe-west9" cluster = google_container_cluster.primary.name node_count = 1 node_config { preemptible = true machine_type = "e2-small" # Google recommends custom service accounts that have cloud-platform scope and permissions granted via IAM Roles. service_account = google_service_account.default.email oauth_scopes = [ "https://www.googleapis.com/auth/cloud-platform" ] } }
核心问题
需要将SSD总用量控制在250GB以内,该如何调整配置?
已尝试操作
- 调整自定义节点池磁盘大小:将自定义节点池的
disk_size_gb设为50GB(默认100GB),但报错信息未变化,修改后的配置:
resource "google_container_node_pool" "primary_preemptible_nodes" { name = "k8s-node-pool" location = "europe-west9" cluster = google_container_cluster.primary.name node_count = 1 node_config { preemptible = true machine_type = "e2-small" disk_size_gb = 50 # Google recommends custom service accounts that have cloud-platform scope and permissions granted via IAM Roles. service_account = google_service_account.default.email oauth_scopes = [ "https://www.googleapis.com/auth/cloud-platform" ] } }
解决方案
问题根源在于临时默认节点池的磁盘大小未被限制:
虽然你设置了remove_default_node_pool = true,但Terraform创建集群时会先初始化默认节点池(initial_node_count = 1),这个节点池默认使用100GB SSD。加上GKE系统组件占用的约100GB SSD,再加上自定义节点池的磁盘容量,总用量会超过250GB配额。
只需在google_container_cluster资源中,给临时默认节点池指定更小的磁盘大小即可,即使它会被删除:
resource "google_container_cluster" "primary" { name = "k8s-cluster" location = "europe-west9" remove_default_node_pool = true initial_node_count = 1 # 给临时默认节点池设置极小磁盘(最小可设10GB) node_config { disk_size_gb = 20 } }
保持自定义节点池disk_size_gb = 50的配置,此时总SSD用量为:系统组件约100GB + 默认节点池20GB + 自定义节点池50GB = 170GB,完全低于250GB配额限制。
内容的提问来源于stack exchange,提问作者itasahobby
相关产品推荐
相关产品推荐

