使用Terraform部署AKS集群时等待完成报错求助
That frustrating 404 error you're hitting is usually either a transient Azure API glitch or a timeout issue. Here are the most effective fixes to get your AKS cluster up and running:
1. Bump Up Terraform's Timeout for AKS Provisioning
Terraform's default timeout for creating an AKS cluster is 30 minutes, but sometimes Azure takes longer to spin up all the underlying resources (especially if your region is under load). Adding an explicit timeout block gives Azure more breathing room:
resource "azurerm_kubernetes_cluster" "k8s" { # ... keep all your existing config here ... timeouts { create = "60m" # Extend to 60 minutes update = "60m" delete = "60m" } }
2. Just Retry the Deployment
More often than not, this error is a temporary hiccup in Azure's API. The deployment might actually be completing in the background, but Terraform can't fetch its status due to a momentary delay. Run terraform apply again—chances are it will succeed on the second try.
3. Double-Check Your Service Principal Permissions
Your service principal needs enough permissions to manage resources in the auto-generated MC_ resource group that AKS creates. Make sure it has at least Contributor access on your subscription (or a custom role that includes permissions to create, read, and update deployments and resource groups).
You can verify this using the Azure CLI:
az role assignment list --assignee <your-sp-client-id> --all
4. Hold Off on Reading the MC Resource Group Early
While your data "azurerm_resource_group" "agents" has a depends_on clause, Terraform's dependency graph might still try to fetch the MC group before it's fully initialized. If you don't need this data source during the initial cluster creation, consider moving it to a separate apply step or module.
5. Check for Regional Azure Issues
Occasionally, this error can stem from regional service outages. Head to the Azure Portal's Service Health section to see if there are any ongoing issues in your target region that might be impacting AKS provisioning.
If all else fails, try destroying the partial deployment with terraform destroy and starting fresh—sometimes leftover resources from a failed attempt can cause weird edge cases.
内容的提问来源于stack exchange,提问作者Anshul Verma

