如何查看Spark处理日志?提交Spark脚本后遇日志问题求助
Hey there! Let's walk through how you can access the detailed Spark logs you need, especially since you're running on YARN and hit that executor-killed error. Here are the most practical methods based on your setup:
1. YARN ResourceManager UI (First Stop for Executor Issues)
Since YARN killed your executor, this is the best place to dig into why it happened:
- Navigate to your YARN ResourceManager web UI (typically at
http://<your-resource-manager-host>:8088) - Head to the Applications tab, then find your Spark app (filter by name or application ID if you have a lot of apps running)
- Click the application ID to open its details page
- Look for the Containers section, find the container linked to your lost executor (executor 3 in your log)
- Click the Logs link for that container—this will show you the executor's stdout and stderr, which should explicitly tell you what memory limits were exceeded
2. Driver Logs (Client Mode Specific)
Since you're using --deploy-mode client, the driver runs on the machine where you ran spark-submit. The logs are already printing to your terminal, but if you want to save them for later review:
Redirect the output when submitting your job to a file:
spark-submit \ --master yarn \ --deploy-mode client \ --driver-memory 2G \ --driver-cores 2 \ --executor-memory 8G \ --num-executors 3 \ --executor-cores 3 \ script.py > spark_driver.log 2>&1
This captures both regular output and errors into spark_driver.log. Also, that warning you saw about spark.yarn.driver.memoryOverhead is normal—it doesn't apply in client mode, so you can ignore that setting for your current deploy mode.
3. Spark History Server
If your cluster has the Spark History Server set up (most do by default), you can access it at http://<history-server-host>:18080:
- Find your running or completed Spark app in the list
- You can view stage-by-stage details, task logs, and memory usage metrics here—super useful for post-job analysis
4. Direct Worker Node Logs
If you have access to the worker node mentioned in your error (serverw19.domain), you can check the local YARN container logs. The default path is usually:
/var/log/hadoop-yarn/containers/
Each container has its own directory named with the application ID and container ID. Inside, you'll find stdout and stderr files with the full executor logs.
Quick note: Since your error is about exceeding memory limits, the executor logs will probably show which part of your job is using more memory than you allocated. You might need to adjust spark.executor.memoryOverhead (separate from --executor-memory) if your job uses off-heap memory.
内容的提问来源于stack exchange,提问作者psmith

