使用CDK部署ECS上KeystoneJS后端遇ALB 502错误求助
问题描述
我用AWS CDK的ApplicationLoadBalancedEc2Service在ECS上部署了KeystoneJS后端,作为AWS新手,访问api.site.com时出现502 Bad Gateway错误。通过ALB访问日志推断错误来自负载均衡器,日志内容如下:
http 2023-03-12T21:26:46.354861Z app/B-Ec2Se-1SJSXLQPN0WT5/060ba3cebce00a91 104.205.179.68:59623 10.0.73.11:49157 -1 -1 -1 502 - 1051 272 "POST http://api.site.com:80/api/graphql HTTP/1.1" "node-fetch" - - arn:aws:elasticloadbalancing:us-east-1:{ACCOUNT_ID}:targetgroup/Ba-Ec2Se-SJ5TSYWCXWKL/4dcb365548b7d993 "Root=1-640e4396-180891915fab0a8a2fdd61a3" "-" "-" 0 2023-03-12T21:26:46.197000Z "forward" "-" "-" "10.0.73.11:49157" "-" "-" "-"
根据AWS文档,该错误由负载均衡器导致,但我看不懂日志中-1 -1 -1对应的错误类型。后端是包含管理面板和GraphQL API的KeystoneJS,怀疑CDK模板配置有误,附上CDK代码:
const vpc = new Vpc(this, "VPC", { cidr: "10.0.0.0/16", subnetConfiguration: [ { name: "elb_public_", subnetType: SubnetType.PUBLIC }, { name: "ecs_private_", subnetType: SubnetType.PRIVATE_WITH_NAT }, { name: "aurora_isolated_", subnetType: SubnetType.PRIVATE_ISOLATED }, ], }); const subnetIds: string[] = []; vpc.isolatedSubnets.forEach((subnet, index) => { subnetIds.push(subnet.subnetId); }); const dbSubnetGroup: CfnDBSubnetGroup = new CfnDBSubnetGroup( this, "AuroraSubnetGroup", { dbSubnetGroupDescription: "Subnet group to access aurora", dbSubnetGroupName: "aurora-serverless-subnet-group", subnetIds, } ); const databaseCredentialsSecret = new secretsManager.Secret( this, "DBCredentialsSecret", { secretName: `${serviceName}-credentials-db`, generateSecretString: { secretStringTemplate: JSON.stringify({ username: databaseUsername, }), excludePunctuation: true, includeSpace: false, generateStringKey: "password", }, } ); new ssm.StringParameter(this, "DBCredentialsArn", { parameterName: `${serviceName}-credentials-arn`, stringValue: databaseCredentialsSecret.secretArn, }); const dbClusterSecurityGroup = new SecurityGroup( this, "DBClusterSecurityGroup", { vpc } ); dbClusterSecurityGroup.addIngressRule( Peer.ipv4("10.0.0.0/16"), Port.tcp(5432) ); const dbConfig = { dbClusterIdentifier: `${serviceName}-cluster`, engineMode: "serverless", engine: "aurora-postgresql", engineVersion: "10.7", databaseName: databaseName, masterUsername: databaseCredentialsSecret .secretValueFromJson("username") .toString(), masterUserPassword: databaseCredentialsSecret .secretValueFromJson("password") .toString(), dbSubnetGroupName: dbSubnetGroup.dbSubnetGroupName, scalingConfiguration: { autoPause: true, maxCapacity: 2, minCapacity: 2, secondsUntilAutoPause: 3600, }, vpcSecurityGroupIds: [dbClusterSecurityGroup.securityGroupId], }; const rdsCluster = new CfnDBCluster(this, "DBCluster", dbConfig); rdsCluster.addDependsOn(dbSubnetGroup); const repo = ecr.Repository.fromRepositoryArn( this, "BackendRepo", `arn:aws:ecr:us-east-1:${ACCOUNT_ID}:repository/backend` ); const image = ecs.ContainerImage.fromEcrRepository(repo); const cluster = new ecs.Cluster(this, "Cluster", { vpc, }); cluster.addCapacity("DefaultAutoScalingGroupCapacity", { instanceType: new ec2.InstanceType("t2.xlarge"), desiredCapacity: 3, }); const loadBalancedService = new ecs_patterns.ApplicationLoadBalancedEc2Service(this, "Ec2Service", { domainName: "api.site.com", memoryLimitMiB: 2048, domainZone: { env: { account: ${ACCOUNT_ID}, region: "us-east-1", }, hostedZoneId: "Z07240521MS2RP8F5K3FW", zoneName: "site.com", hostedZoneArn: "arn:aws:route53:::hostedzone/Z07240521MS2RP8F5K3FW", stack: this, node: this.node, applyRemovalPolicy: () => RemovalPolicy.DESTROY, }, cluster, taskImageOptions: { image: image, environment: { ... }, }, });
问题分析与解决
1. ALB日志中-1 -1 -1的含义
这三个-1对应ALB日志字段里的target_processing_time、response_processing_time、elb_status_code,全部为-1说明:
- 负载均衡器尝试连接目标实例(ECS容器),但没有收到任何响应,连接可能被拒绝或超时。
2. 最可能的配置问题:容器端口不匹配
KeystoneJS默认启动在3000端口,但ApplicationLoadBalancedEc2Service的默认容器端口是80。你的CDK代码里没有指定containerPort,导致ALB把请求转发到容器的80端口,而KeystoneJS根本没在这个端口监听,直接引发连接失败,返回502。
修改方法:在taskImageOptions里添加containerPort配置:
taskImageOptions: { image: image, environment: { // 你的环境变量 }, containerPort: 3000 // 添加这一行,指定KeystoneJS的监听端口 }
3. 其他排查点
- 安全组配置:检查ECS任务的安全组是否允许ALB的安全组访问3000端口。CDK默认会创建安全组,但如果有自定义安全组,需确保Ingress规则允许ALB的IP或安全组访问3000端口。
- 容器启动状态:查看ECS任务的日志,确认KeystoneJS是否正常启动,有没有数据库连接失败、环境变量缺失等错误。如果容器启动失败,目标组会把实例标记为不健康,ALB也会返回502。
- 目标组健康检查:默认的健康检查路径是
/,如果KeystoneJS的根路径没有返回200状态码(比如管理面板需要登录),需修改健康检查路径为KeystoneJS的健康端点,比如/api/health(如果你的KeystoneJS配置了的话)。可以通过CDK的targetGroup配置修改:
const loadBalancedService = new ecs_patterns.ApplicationLoadBalancedEc2Service(this, "Ec2Service", { // 其他配置 targetGroup: { healthCheck: { path: "/api/health", port: "3000" } } });
内容的提问来源于stack exchange,提问作者FRMR
相关产品推荐
相关产品推荐

