如何获取Spark应用的CPU、内存使用率及平均资源利用率用于监控
Hey there! Let's walk through how to grab memory and CPU usage for your Spark applications, plus set up polling to calculate average resource utilization for monitoring needs.
There are several reliable ways to get these metrics, depending on your use case (debugging vs. production monitoring):
1. Spark UI (最直观的调试方式)
Spark's built-in UI is the easiest place to start:
- Fire up the UI at
http://<driver-host>:4040(default port; use 18080 for the History Server if the app has finished) - Navigate to the Executors tab: You'll see a breakdown of each executor's used memory (
Used Memory), allocated cores, and even CPU utilization (some versions showCPU Timewhich you can compare to wall-clock time to calculate usage) - The Application Summary section on the homepage also shows total allocated vs. used resources for the entire app.
2. Spark REST API (适合自动化脚本)
The Spark UI is powered by a REST API, so you can programmatically pull metrics:
- Use the executors endpoint to get per-executor resource data:
curl http://<driver-host>:4040/api/v1/applications/<your-app-id>/executors - The JSON response includes key fields like:
coresUsed: Number of CPU cores currently in use by the executormemoryUsed: Used memory in bytestotalCores: Total cores allocated to the executormaxMemory: Maximum memory allocated to the executor
3. Spark Metrics System (生产级监控)
Spark has a built-in metrics system that can export data to tools like JMX, Graphite, or Prometheus:
- Configure metrics by editing
conf/metrics.properties(copy the templatemetrics.properties.templatefirst) - For example, to enable JMX metrics for executors:
executor.source.jvm.class=org.apache.spark.metrics.source.JvmSource executor.sink.jmx.class=org.apache.spark.metrics.sink.JmxSink - You can then use tools like
jconsoleorjvisualvmto connect to the executor/driver JVMs and access metrics like:executor.process.cpu.usage: CPU usage percentageexecutor.jvm.total.used: Used heap memory
4. 命令行工具 (单节点调试)
If you just need quick stats for a running driver/executor process on a single node:
- Use
psto find Spark processes and check CPU/memory:ps aux | grep spark - Look at the
%CPUcolumn for CPU usage andRSS(Resident Set Size) for memory usage (convert to MB by dividing by 1024).
To track average utilization over time, you'll need to set up periodic polling and aggregate the collected data:
方法1: 自定义Shell/Python脚本
A simple script can call the Spark REST API on a schedule, store metrics, and calculate averages later:
Example Shell Script (with curl and jq)
#!/bin/bash # 配置参数 APP_ID="your-spark-application-id" DRIVER_URL="http://localhost:4040" POLL_INTERVAL=60 # 每60秒轮询一次 OUTPUT_FILE="spark_resource_metrics.csv" # 初始化CSV表头 echo "timestamp,total_cores_used,total_memory_used_mb" > $OUTPUT_FILE echo "Starting resource polling... Press Ctrl+C to stop." while true; do # 获取当前时间戳 TIMESTAMP=$(date +%Y-%m-%dT%H:%M:%S) # 调用REST API获取executors数据 RESPONSE=$(curl -s "$DRIVER_URL/api/v1/applications/$APP_ID/executors") # 计算总CPU核心使用数 TOTAL_CORES=$(echo $RESPONSE | jq '[.[] | .coresUsed] | add') # 计算总内存使用量(转换为MB) TOTAL_MEM_MB=$(echo $RESPONSE | jq '[.[] | .memoryUsed] | add / 1024 / 1024') # 将数据写入CSV echo "$TIMESTAMP,$TOTAL_CORES,$TOTAL_MEM_MB" >> $OUTPUT_FILE sleep $POLL_INTERVAL done
计算平均利用率
Once you've collected data, use awk to compute averages:
awk -F ',' 'NR>1 {cpu_sum+=$2; mem_sum+=$3; count++} END { print "Average CPU Cores Used: " cpu_sum/count print "Average Memory Used (MB): " mem_sum/count }' spark_resource_metrics.csv
方法2: 集成监控系统 (生产环境推荐)
For long-term monitoring, integrate Spark with tools like Prometheus + Grafana:
- Configure Spark to export metrics to Prometheus:
Add these settings tospark-defaults.conf:spark.metrics.conf.driver.sink.prometheus.class=org.apache.spark.metrics.sink.PrometheusSink spark.metrics.conf.executor.sink.prometheus.class=org.apache.spark.metrics.sink.PrometheusSink spark.metrics.conf.sink.prometheus.port=9104 - Set up Prometheus to scrape the Spark metrics endpoint.
- Build dashboards in Grafana:
- Create panels to visualize CPU/memory usage over time
- Use Grafana's built-in functions to calculate average utilization for any time range
- Set up alerts if utilization exceeds thresholds
内容的提问来源于stack exchange,提问作者Vishal

