如何将Spark Streaming输入行数、批处理时长等指标接入Prometheus?
如何将Spark Streaming流处理指标接入Prometheus
问题背景
已知Spark 3.0支持将Streaming的输入行数、处理速率、批处理时长等指标接入Prometheus,已配置Spark集群开启相关监控(spark.ui.prometheus.enabled true、spark.sql.streaming.metricsEnabled true),并能通过指定API获取主机/基础设施指标,但无法找到获取流处理相关指标的API或参考资料,需解决如何将流UI指标接入Prometheus。
解决方案
1. 切换到Driver端的Prometheus指标端点
流处理的核心指标属于Spark Driver的监控范畴,而非Executor。你需要将Prometheus的抓取路径从Executor端点切换到Driver的Prometheus指标端点:
原使用的Executor指标路径:
/driver-proxy-api/o/<org-id>/<cluster-id>/40001/metrics/executors/prometheus
替换为Driver指标路径:
/driver-proxy-api/o/<orgid>/<clusterid>/40001/metrics/prometheus
该端点会返回Driver侧的所有监控指标,包括流处理相关的spark_sql_streaming_前缀系列指标(如spark_sql_streaming_input_rows_total、spark_sql_streaming_batch_duration_seconds等)。
2. 更新Prometheus配置
建议新增独立的抓取任务专门采集流处理指标,避免与基础设施指标混淆。修改后的Prometheus配置示例如下:
# my global config global: scrape_interval: 15s # 抓取间隔设为15秒,默认1分钟 evaluation_interval: 15s # 规则评估间隔设为15秒,默认1分钟 # scrape_timeout 使用全局默认值(10秒) # Alertmanager配置 alerting: alertmanagers: - static_configs: - targets: # - alertmanager:9093 # 加载规则并按全局evaluation_interval定期评估 rule_files: # - "first_rules.yml" # - "second_rules.yml" # 抓取配置 scrape_configs: # 采集主机/Executor基础设施指标 - job_name: 'spark_executors_infra' scheme: https scrape_interval: 5s static_configs: - targets: ['eastus-c3.azuredatabricks.net'] metrics_path: '/driver-proxy-api/o/<orgid>/<clusterid>/40001/metrics/executors/prometheus' basic_auth: username: 'token' password: '<你的用户生成token>' # 采集Spark流处理指标 - job_name: 'spark_streaming_metrics' scheme: https scrape_interval: 5s static_configs: - targets: ['eastus-c3.azuredatabricks.net'] metrics_path: '/driver-proxy-api/o/<orgid>/<clusterid>/40001/metrics/prometheus' basic_auth: username: 'token' password: '<你的用户生成token>'
3. 验证与生效
- 用浏览器或curl工具访问修改后的Driver指标端点,确认返回结果中包含
spark_sql_streaming_前缀的指标(需确保流作业处于运行状态,未运行的流作业不会生成相关指标)。 - 重启Prometheus服务使配置生效,之后即可在Prometheus UI中查询流处理相关指标。
注意事项
- 确保你的用户token拥有集群的查看权限,否则无法访问driver-proxy-api端点。
- 若集群为自动终止模式,需保证集群处于运行状态才能正常抓取指标。
内容的提问来源于stack exchange,提问作者Shane
相关产品推荐
相关产品推荐

