Spark History Server查看已完成YARN应用容器日志时重定向失败
问题描述
Spark版本3.3.0,Hadoop版本3.3.1,在Spark History Server查看已完成的YARN集群Spark应用日志时,点击应用「Executors」页面的stdout或stderr链接,未显示预期的executor容器日志,而是出现「Failed redirect」页面。
已配置的相关参数如下:
spark-defaults.conf 配置
## Spark History Server spark.yarn.historyServer.address my-spark3-history-server.com spark.eventLog.enabled true spark.eventLog.compress false spark.eventLog.dir s3a://TMP_BUCKET/PREFIX/spark3/application-history spark.history.fs.logDirectory s3a://TMP_BUCKET/PREFIX/spark3/application-history spark.history.fs.cleaner.enabled true spark.history.fs.cleaner.maxAge 7d spark.history.fs.inProgressOptimization.enabled true spark.history.store.maxDiskUsage 200g spark.history.store.path /<some-dir>/spark-history ## Uncommenting below line just makes the log URL redirect to Spark-History server's main page # spark.history.custom.executor.log.url https://my-spark3-history-server.com/jobhistory/logs/{{NM_HOST}}:{{NM_PORT}}/{{CONTAINER_ID}}/{{CONTAINER_ID}}/{{USER}}/{{FILE_NAME}}?start=-4096
yarn-site.xml 配置
<property> <name>yarn.log-aggregation-enable</name> <value>true</value> <description> Whether to enable log aggregation. Log aggregation collects each container's logs and moves these logs onto a file-system, for e.g. HDFS, after the application completes. Users can configure the "yarn.nodemanager.remote-app-log-dir" and "yarn.nodemanager.remote-app-log-dir-suffix" properties to determine where these logs are moved to. Users can access the logs via the Application Timeline Server. </description> </property> <property> <name>yarn.log.server.url</name> <value>https://my-spark3-history-server.com/jobhistory/logs/</value> <description>URL for log aggregation server</description> </property> <property> <name>yarn.nodemanager.log-aggregation.roll-monitoring-interval-seconds</name> <value>1800</value> </property> <property> <name>yarn.nodemanager.remote-app-log-dir</name> <value>s3a://BUCKET_NAME/PREFIX_NAME/yarn3_logs/</value> <description>Where to aggregate logs to.</description> </property> <property> <name>yarn.log-aggregation.retain-seconds</name> <value>604800</value> <description> How long to keep aggregation logs before deleting them. -1 disables. Be careful set this too small and you will spam the name node. We set to 7 days. </description> </property> <property> <name>yarn.log-aggregation.retain-check-interval-seconds</name> <value>-1</value> <description> How long to wait between aggregated log retention checks. If set to 0 or a negative value then the value is computed as one-tenth of the aggregated log retention time. Be careful set this too small and you will spam the name node. </description> </property> <property> <name>yarn.nodemanager.local-dirs</name> <value>{{ yarn_local_dirs }}</value> </property>
已确认S3上spark.eventLog.dir和yarn.nodemanager.remote-app-log-dir对应的目录均有数据。
疑问:即使YARN仅运行Spark任务,是否也需要同时运行Spark历史服务器和MapReduce历史服务器?
解决方案与说明
关于是否需要MapReduce历史服务器
是的,即使YARN仅运行Spark任务,也需要同时运行Spark History Server和MapReduce History Server。原因如下:
- Spark History Server仅负责解析Spark事件日志、生成应用执行历史页面,本身不具备直接读取和展示YARN聚合日志的能力。
- YARN容器日志聚合后,需通过MapReduce History Server(Hadoop 3.x中可结合YARN Application Timeline Server)提供的服务访问,Spark History Server生成的
stdout/stderr链接本质是跳转到MapReduce History Server的日志端点。
针对当前配置的排查与修复建议
确认MapReduce History Server已启动
检查集群中MapReduce History Server的运行状态,确保服务正常启动,且yarn.log.server.url配置的地址正确指向该服务的日志路径(当前配置的https://my-spark3-history-server.com/jobhistory/logs/应为MapReduce History Server端点,而非Spark History Server)。调整自定义日志链接模板(可选)
若启用spark.history.custom.executor.log.url,需确保模板中的变量({{NM_HOST}},{{NM_PORT}},{{CONTAINER_ID}},{{USER}},{{FILE_NAME}})能被正确替换,且路径匹配MapReduce History Server的日志访问规则。之前启用后跳转到Spark History主页,说明模板路径错误,可调整为类似:spark.history.custom.executor.log.url=https://my-mapreduce-history-server.com/jobhistory/logs/{{NM_HOST}}:{{NM_PORT}}/{{CONTAINER_ID}}/{{USER}}/{{FILE_NAME}}?start=-4096验证日志聚合的完整性
检查S3中yarn.nodemanager.remote-app-log-dir下的日志结构,确认每个Spark应用的容器日志已正确归档到user/<用户名>/logs/<应用ID>/路径下,且MapReduce History Server有权限读取这些S3文件。排查服务日志
查看Spark History Server和MapReduce History Server的日志,查找与日志链接生成、日志访问相关的错误信息,进一步定位问题。
内容的提问来源于stack exchange,提问作者akki

