作为Hadoop普通用户,无CLI权限如何查看运行中作业详情?
Since you don't have CLI access but can use Hue and Jupyter, here are practical, actionable ways to get the detailed job info you need:
1. Query Spark Status Directly via Jupyter Notebook
You can bypass the broken UI entirely by using Spark's built-in status tracker API right in your Jupyter PySpark (or Scala) notebook. This gives you real-time, granular details about your jobs.
Try running these code snippets in your notebook:
# Fetch basic application configuration details app_config = spark.sparkContext.getConf().getAll() print("Application Configuration:") for key, val in app_config: print(f" {key}: {val}") # Get all active jobs with core status info active_jobs = spark.sparkContext.statusTracker().getAllJobs() print("\nActive Spark Jobs:") for job in active_jobs: print(f" Job ID: {job.jobId()}") print(f" Status: {job.status()}") print(f" Submission Time: {job.submissionTime()}\n") # Drill into stage-level details for a specific job if active_jobs: target_job_id = active_jobs[0].jobId() stage_ids = spark.sparkContext.statusTracker().getJobInfo(target_job_id).stageIds() print(f"Stage Details for Job {target_job_id}:") for stage_id in stage_ids: stage_info = spark.sparkContext.statusTracker().getStageInfo(stage_id) print(f" Stage ID: {stage_id}") print(f" Stage Name: {stage_info.name()}") print(f" Tasks: {stage_info.numCompletedTasks()}/{stage_info.numTasks()} completed") print(f" Duration: {stage_info.duration()/1000 if stage_info.duration() else 'N/A'} seconds\n")
This will pull job status, stage progress, task completion rates, and application configs directly from the Spark context you're already using—no external UI required.
2. Use Hue's Integrated Job Monitoring Features
Even if the native Hadoop UI links are broken, Hue has its own tools to track jobs:
- Hue Job Browser: Navigate to the Job Browser, filter jobs by your username (since you submitted them) and status (Running/Submitted). Here you can view the job's DAG visualization, execution timeline, and basic resource usage stats.
- Hive Query History: If your Spark jobs are triggered by Hive queries, check Hue's Hive query history. Each query links to its corresponding Spark job, and Hue often retains logs and execution plan details even when the native UI fails.
- YARN Cluster View: Look for the Clusters tab in Hue (if enabled). You can access the YARN application list directly through Hue, filter by your user, and check the Logs tab—sometimes Hue can pull container logs even when the attempt details page throws an error.
3. Request Configuration Fixes from Your Ops Team
The "Failed to read the attempts of the application" error is almost always a permissions or UI setup issue:
- Your user might not have access to the ResourceManager API endpoints that fetch attempt details.
- The Hadoop UI proxy configuration could be broken for your user role.
Reach out to your ops team with specific context: mention that you can't view attempt details for jobs you submitted, even though the UI lists the applications. Ask them to either grant you the necessary API permissions or fix the UI proxy setup—this will let you use the native Hadoop UI as intended long-term.
内容的提问来源于stack exchange,提问作者TheProletariat

