AWS Batch无法创建c5d.18xlarge实例,提交对应作业久未启动
Let’s walk through the most likely causes and actionable fixes for your stuck job and instance creation issue:
1. Check Account Service Quotas for c5d.18xlarge
First, verify if your AWS account has sufficient quota for c5d.18xlarge instances in your region:
- Head to the AWS Service Quotas console, search for "EC2 On-Demand Instances" (or "EC2 Spot Instances" if you’re using Spot capacity)
- Locate the quota entry for
c5d.18xlarge—if your current usage hits the limit, AWS can’t provision additional instances. You’ll need to submit a quota increase request directly via the console.
2. Validate Compute Environment Instance Configuration
Ensure your AWS Batch compute environment is set to allow c5d.18xlarge instances:
- For managed compute environments, check the instance types specified in the environment settings. It should either explicitly include
c5d.18xlargeor cover it via a family wildcard likec5d(which includes all sizes in the c5d family). If missing, update the compute environment to add this instance type. - For unmanaged environments, confirm your self-managed instances are of the
c5d.18xlargetype and properly registered with Batch.
3. Verify Job Queue & Compute Environment Association
Double-check that your job queue is linked to the correct compute environment (the one with a 256 max vCPU limit):
- In the AWS Batch console, navigate to your job queue and review the list of associated compute environments. If the wrong environment is tied to the queue, your job won’t be routed to the right resource pool.
4. Check Regional Instance Capacity
Even with sufficient quota, AWS may have limited on-demand/spot capacity for c5d.18xlarge in your region. Test this by:
- Manually launching a standalone
c5d.18xlargeEC2 instance via the EC2 console. If this fails with a capacity error, you have a few options:- Switch to an alternative instance type (e.g.,
c5.18xlargeor another high-CPU family that has available capacity) - Request a capacity reservation for
c5d.18xlarge, or wait for regional capacity to replenish.
- Switch to an alternative instance type (e.g.,
5. Validate Container Override Settings
Your container configuration looks mostly correct, but let’s confirm a few details:
- The memory value
144000uses MiB (AWS Batch’s default unit for container memory). Sincec5d.18xlargehas 147456 MiB of total memory, this value is well within limits—just ensure you didn’t accidentally specify GiB (which would be drastically over the instance’s capacity). - The vCPU count
72matches the instance’s exact vCPU count. AWS Batch requires container vCPUs to align with the instance’s available resources (minus minimal overhead for the Batch agent), so this configuration is valid.
6. Dig Into Job Events & Logs
For more granular troubleshooting, check the job’s event stream and related logs:
- In the AWS Batch console, select your stuck job and view the "Events" tab. This will show specific messages like "Job waiting for capacity" or "Instance creation failed" with error codes that point to the root cause.
- Check CloudWatch logs for the Batch compute environment’s EC2 instances (if any were attempted to launch) to spot errors during instance initialization.
内容的提问来源于stack exchange,提问作者Pyae Phyoe Shein

