You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Prometheus监控Docker容器timesheet的运行状态?

用Prometheus和cAdvisor判断Docker容器(timesheet)的运行状态

一、直接查询指标判断容器在线/离线

你可以直接在Prometheus的表达式浏览器中执行以下PromQL语句,快速判断timesheet容器的状态:

  • 判断容器是否正在运行
    用container_state_running指标(gauge类型,1=运行中,0=未运行):

    container_state_running{name="timesheet"}
    

    返回值为1 → 容器在线;无返回结果或值为0 → 容器已停止/离线。

  • 判断容器是否长时间未被采集
    用container_last_seen指标记录cAdvisor最后一次采集到容器数据的时间戳,和当前时间对比:

    time() - container_last_seen{name="timesheet"} > 30
    

    如果结果为true,说明容器超过30秒没被采集到(可根据cAdvisor采集间隔调整阈值),大概率已离线。

  • 验证容器是否曾启动过
    查询container_start_time_seconds指标:

    container_start_time_seconds{name="timesheet"}
    

    能查到结果 → 容器至少启动过;指标消失 → 容器已被销毁或彻底离线。

二、配置Prometheus实现容器存活监控与告警

1. 添加告警规则

在Prometheus的告警规则文件(比如alerting_rules.yml)中添加针对timesheet容器的离线告警:

groups:
- name: container_alerts
  rules:
  - alert: TimesheetContainerDown
    expr: container_state_running{name="timesheet"} == 0 OR absent(container_state_running{name="timesheet"})
    for: 1m
    labels:
      severity: critical
    annotations:
      summary: "Timesheet容器已离线"
      description: "容器timesheet停止运行或无法被cAdvisor采集,持续时间超过1分钟"

添加后重启Prometheus加载规则,当容器离线超过1分钟时,Prometheus会触发告警。

2. 结合Docker健康检查增强监控(可选)

如果容器支持健康检查,可以先给timesheet配置Docker健康检查:

# 示例:用curl检查容器内的健康接口,可根据实际调整
docker run --name timesheet \
  --health-cmd="curl -f http://localhost/health || exit 1" \
  --health-interval=30s \
  --health-timeout=5s \
  your-timesheet-image:tag

配置后,cAdvisor会采集container_health_status指标,用以下PromQL判断健康状态:

container_health_status{name="timesheet"} == 0

值为1 → 容器健康运行;0 → 容器不健康;无结果 → 容器离线。

3. 可视化监控

如果使用Grafana,搜索并导入cAdvisor相关仪表盘,添加Prometheus作为数据源后,筛选timesheet容器的状态面板,即可实时查看容器的运行状态、资源占用等信息。

内容的提问来源于stack exchange,提问作者Andrew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 04:57:09