AWS Elastic Beanstalk部署Docker版Prometheus/Grafana无法访问求助
我尝试在AWS Elastic Beanstalk的Docker环境中部署Prometheus和Grafana监控栈,本地测试一切正常,但部署到云端后遇到了严重的健康问题,且完全无法访问服务。以下是详细情况:
环境配置
Dockerrun.aws.json 配置
{ "AWSEBDockerrunVersion": "2", "volumes": [ { "name": "prometheus-conf", "host": { "sourcePath": "/var/app/current/prometheus" } } ], "containerDefinitions": [ { "name": "prometheus-app", "image": "prom/prometheus", "essential": true, "memory": 512, "portMappings": [ { "hostPort": 9090, "containerPort": 9090 } ], "mountPoints": [ { "sourceVolume": "prometheus-conf", "containerPath": "/opt/prometheus" } ], "command": [ "--config.file=/opt/prometheus/prometheus.yml" ] }, { "name": "grafana", "image": "grafana/grafana", "essential": true, "memory": 256, "links": [ "prometheus-app" ], "portMappings": [ { "hostPort": 3000, "containerPort": 3000 } ] } ] }
Prometheus 配置 (prometheus/prometheus.yml)
global: scrape_interval: 15s evaluation_interval: 15s rule_files: scrape_configs: - job_name: 'prometheus' static_configs: - targets: ['localhost:9090'] - job_name: 'jParser' static_configs: - targets: ['jParser.fpemryt2er.us-east-2.elasticbeanstalk.com']
本地测试情况
使用 eb cli 执行 eb local run 后,本地运行完全正常:
- 可通过
localhost:9090正常访问Prometheus - 可通过
localhost:3000正常访问Grafana
部署后问题
执行 eb deploy 部署到Elastic Beanstalk后:
- 部署流程显示成功,EC2控制台可见实例处于运行状态
- 仅1分钟后,环境状态从OK直接转为Severe,收到系统提示:
Environment health has transitioned from Ok to Severe. ELB health is failing or not available for all instances.
100.0 % of the requests to the ELB are failing with HTTP 5xx (10 minutes ago) - 尝试多种方式均无法访问服务:
- Elastic Beanstalk URL + 端口(如
xxx.elasticbeanstalk.com:9090) - EC2实例公网DNS/IP + 端口
- Elastic Beanstalk URL + 端口(如
已排查的操作
- 已在负载均衡器的安全组中开放了9090和3000端口
可能的原因与解决步骤
本地正常但云端出问题,大概率是Elastic Beanstalk的健康检查配置、容器启动异常或者网络权限细节没处理好。我给你梳理几个关键排查方向:
1. 修正Elastic Beanstalk健康检查配置
Elastic Beanstalk默认的健康检查是针对80端口的/路径,但你的服务运行在9090和3000端口,这会导致ELB误判实例不健康,进而触发5xx错误。
解决方法:
- 进入Elastic Beanstalk控制台,找到你的环境 → Configuration → Load balancer
- 在Health check部分,修改:
- 端口:选择
9090(优先监控核心的Prometheus服务) - 路径:改为
/-/healthy(Prometheus官方提供的健康检查端点)
- 端口:选择
- 保存配置后等待环境更新,观察健康状态是否恢复
2. 验证全链路安全组权限
虽然你开放了ELB的端口,但还要确认两个关键点:
- EC2实例的安全组:是否允许来自ELB安全组的9090、3000端口流量,以及ELB的健康检查流量(通常是TCP或HTTP协议)
- ELB的安全组:是否允许外部(0.0.0.0/0)访问9090和3000端口
- 额外检查VPC的网络ACL,确认没有阻断这些端口的入站/出站规则
3. 检查容器启动日志与运行状态
EC2实例看起来正常,但容器可能启动失败或有隐性错误:
- 通过EB控制台的Logs选项卡,下载完整日志包,重点查看
docker相关的启动日志 - 或者用EB CLI执行:
eb logs,快速拉取最新的环境日志 - 直接登录EC2实例,执行
docker ps查看容器是否在运行;如果容器已退出,用docker logs <container-id>查看具体错误(比如Prometheus是否找不到配置文件)
4. 确认Prometheus配置文件挂载有效性
你的Dockerrun配置中挂载了本地prometheus目录到容器,但要验证:
- 部署后EC2实例的
/var/app/current/prometheus目录是否存在,prometheus.yml文件是否正确上传 - 容器内的
/opt/prometheus/prometheus.yml是否有正确的读取权限(Prometheus进程需要能访问该文件) - 可以登录EC2实例,执行
docker exec -it <prometheus-container-id> ls /opt/prometheus验证配置文件是否存在
5. 排查Grafana的容器间通信
虽然配置了links到prometheus-app,但在Elastic Beanstalk的Docker环境中,容器间通信建议直接用容器名称作为地址。等你能访问Grafana后,配置数据源时可以尝试用prometheus-app:9090作为目标地址。
优先解决Prometheus的启动和健康检查问题,再处理Grafana的访问配置会更高效。
内容的提问来源于stack exchange,提问作者Oleg Shankovskyi

