GCE中Python prometheus_client向Grafana Cloud上报指标异常求助
问题背景
作为Grafana Cloud/Prometheus新手,尝试将Google Compute实例中Python prometheus_client生成的指标上报到Grafana Cloud,但初期找不到上报的指标。后来发现需要配置static_configs targets,同时在Python脚本中调用start_http_server并创建Summary类型指标后,Agent才开始向云端上报指标,但不清楚其中原因。
Python代码
from prometheus_client import Summary, CollectorRegistry, start_http_server from time import perf_counter c_registry = CollectorRegistry() api_hits_summary = Summary('resp_time','API calls', ['endpoint'], registry=c_registry) start_http_server(8000) ... st = perf_counter() _, raw_resp = self._h.request(url) api_hits_summary.labels(endpoint='S').observe(perf_counter()-st)
排查细节
- 执行
curl localhost:8000后,未看到resp_time_created相关内容 - 确认Agent指标已在云端数据源中,
prometheus_wal_watcher_current_segment与prometheus_tsdb_wal_segment_current指标完全匹配 - 仅看到少量警告日志:
May 07 11:49:53 instance-2 grafana-agent[17804]: ts=2023-05-07T15:49:53.111085331Z caller=wal.go:409 level=info agent=prometheus instance=<I removed instance-id> msg="series GC completed" duration=2.816839ms May 07 11:49:53 instance-2 grafana-agent[17804]: ts=2023-05-07T15:49:53.112838837Z caller=checkpoint.go:100 level=info agent=prometheus instance=<I removed instance-id> msg="Creating checkpoint" from_segment=46 to_segment=49 mint=1683474270000 May 07 11:49:53 instance-2 grafana-agent[17804]: ts=2023-05-07T15:49:53.146233662Z caller=cleaner.go:203 level=warn agent=prometheus component=cleaner msg="unable to find segment mtime of WAL" name=/var/lib/grafana-agent/.cache err="unable to open WAL: open /var/lib/grafana-agent/.cache/wal: no such file or directory" May 07 11:49:53 instance-2 grafana-agent[17804]: ts=2023-05-07T15:49:53.926131288Z caller=wal.go:474 level=info agent=prometheus instance=<I removed instance-id> msg="WAL checkpoint complete" first=46 last=49 duration=817.8648ms May 07 12:19:53 instance-2 grafana-agent[17804]: ts=2023-05-07T16:19:53.248281806Z caller=cleaner.go:203 level=warn agent=prometheus component=cleaner msg="unable to find segment mtime of WAL" name=/var/lib/grafana-agent/.cache err="unable to open WAL: open /var/lib/grafana-agent/.cache/wal: no such file or directory" May 07 12:49:53 instance-2 grafana-agent[17804]: ts=2023-05-07T16:49:53.036460234Z caller=cleaner.go:203 level=warn agent=prometheus component=cleaner msg="unable to find segment mtime of WAL" name=/var/lib/grafana-agent/.cache err="unable to open WAL: open /var/lib/grafana-agent/.cache/wal: no such file or directory" May 07 12:49:54 instance-2 grafana-agent[17804]: ts=2023-05-07T16:49:54.094024836Z caller=wal.go:409 level=info agent=prometheus instance=<I removed instance-id> msg="series GC completed" duration=106.864899ms May 07 12:49:54 instance-2 grafana-agent[17804]: ts=2023-05-07T16:49:54.21561055Z caller=checkpoint.go:100 level=info agent=prometheus instance=<I removed instance-id> msg="Creating checkpoint" from_segment=50 to_segment=51 mint=1683477870000 May 07 12:49:54 instance-2 grafana-agent[17804]: ts=2023-05-07T16:49:54.469656064Z caller=wal.go:474 level=info agent=prometheus instance=<I removed instance-id> msg="WAL checkpoint complete" first=50 last=51 duration=482.495692ms May 07 13:19:53 instance-2 grafana-agent[17804]: ts=2023-05-07T17:19:53.151303299Z caller=cleaner.go:203 level=warn agent=prometheus component=cleaner msg="unable to find segment mtime of WAL" name=/var/lib/grafana-agent/.cache err="unable to open WAL: open /var/lib/grafana-agent/.cache/wal: no such file or directory"
- WAL目录存在文件:
instance-2:~$ sudo ls -l /var/lib/grafana-agent/<I removed instance-id>/wal/ total 864K -rw-r--r-- 1 grafana-agent grafana-agent 288K May 7 11:49 00000052 -rw-r--r-- 1 grafana-agent grafana-agent 288K May 7 12:49 00000053 -rw-r--r-- 1 grafana-agent grafana-agent 271K May 7 13:47 00000054 drwxr-xr-x 2 grafana-agent grafana-agent 4.0K May 7 12:49 checkpoint.00000051
当前配置文件(/etc/grafana-agent.yaml)
server: log_level: info metrics: global: scrape_interval: 1m remote_write: - url: https://prometheus-prod-<url>.grafana.net/api/prom/push basic_auth: username: <userid> password: <api key> wal_directory: '/var/lib/grafana-agent' configs: # Example Prometheus scrape configuration to scrape the agent itself for metrics. # This is not needed if the agent integration is enabled. # - name: agent # host_filter: false # scrape_configs: # - job_name: agent # static_configs: # - targets: ['127.0.0.1:9090'] integrations: agent: enabled: true node_exporter: enabled: true include_exporter_metrics: true disable_collectors: - "mdadm"
问题解答
1. 为什么需要配置static_configs targets?
Grafana Agent是Prometheus的轻量实现,遵循主动拉取指标的工作模式。你必须在Agent的scrape配置中指定拉取目标(即你的Python服务localhost:8000),否则Agent不知道从哪里获取自定义指标。当前配置缺少针对Python服务的scrape任务,需在metrics.configs中新增:
metrics: configs: - name: python-app host_filter: false scrape_configs: - job_name: python-api-metrics static_configs: - targets: ['127.0.0.1:8000']
2. 为什么调用start_http_server和创建Summary指标后才开始上报?
start_http_server(8000):启动HTTP服务器暴露Prometheus格式的指标端点,这是Agent拉取指标的唯一入口,没有这个服务器Agent无法访问你的指标。- 创建
Summary并调用observe:Prometheus客户端仅在指标有实际数据写入后,才会将其暴露到HTTP端点。如果没有执行observe写入数据,端点不会出现resp_time相关指标,Agent自然拉取不到。
3. 为什么curl localhost:8000看不到resp_time_created?
你的代码存在关键问题:start_http_server默认使用全局注册表,但你的resp_time指标注册在自定义的c_registry里,导致HTTP端点根本不会暴露该指标。需要修改启动代码,关联自定义注册表:
start_http_server(8000, registry=c_registry)
此外,resp_time_created是Summary的辅助指标,只有当resp_time有数据写入后才会生成,需确保api_hits_summary.labels(...).observe(...)被执行过。
4. WAL警告日志的影响?
日志中unable to open WAL: open /var/lib/grafana-agent/.cache/wal: no such file or directory是因为Agent配置的wal_directory为/var/lib/grafana-agent,但实际WAL文件存储在实例ID子目录下,该警告不影响指标上报,属于内部缓存目录问题。可手动创建目录解决:
sudo mkdir -p /var/lib/grafana-agent/.cache/wal sudo chown grafana-agent:grafana-agent /var/lib/grafana-agent/.cache/wal
内容的提问来源于stack exchange,提问作者knotmine

