Hadoop集群YARN Dr.who Application Attempt失败,两千+应用无因崩溃求助
Hey there, sorry to hear you're stuck with this massive wave of failing YARN app attempts—2000+ is a tough spot to be in! Let me walk through the most common root causes behind these mysterious Dr.who failures, based on years of troubleshooting Hadoop clusters:
1. Cluster Resource Starvation
Dr.who is YARN's "fallback" placeholder application that gets spawned when the cluster can't allocate critical resources (memory, CPU, disk) to your actual jobs. If thousands of these are failing, your cluster is almost certainly saturated. Here's how to verify:
- Run
yarn node -listto check if all nodes are reporting near-max resource usage - Pull up the YARN ResourceManager UI's Cluster Metrics tab to see remaining memory/vcores
- Hunt for stuck, long-running jobs that are hogging resources and haven't been cleaned up
2. Misconfigured YARN Scheduler Settings
Incorrect scheduler rules can block resource allocation even if some capacity is available. Common issues include:
- Capacity Scheduler: Check if your target queue has hit its maximum capacity or has strict limits on application attempts per queue
- Fair Scheduler: Ensure there's no misconfigured minimum share that's locking out new applications
- Verify settings like
yarn.scheduler.maximum-allocation-mboryarn.scheduler.maximum-allocation-vcoresaren't set too low for your job requirements
3. Unhealthy NodeManagers
If a large chunk of NodeManagers are unresponsive or unhealthy, the ResourceManager can't find valid nodes to launch app attempts. Dig into these checks:
- Check NodeManager logs (typically in
$HADOOP_HOME/logs/yarn-nodemanager-*.log) for errors like disk failures, memory leaks, or lost connections to the ResourceManager - Use
yarn node -status <node-id>to inspect individual node health statuses - Confirm NodeManagers have enough free disk space—YARN marks nodes unhealthy if usage exceeds the threshold set in
yarn.nodemanager.disk-health-checker.max-disk-usage-percentage
4. ResourceManager State Corruption
While rare, a corrupted internal state in the ResourceManager can cause it to incorrectly reject valid application attempts. Try these steps (note: plan for downtime, as this will interrupt running jobs):
- Restart the ResourceManager service
- Scan the ResourceManager logs (
$HADOOP_HOME/logs/yarn-resourcemanager-*.log) for unusual errors triggered during app submission
5. Network/DNS Communication Failures
Silent communication drops between clients, ResourceManager, and NodeManagers can lead to unexplained failures. Verify:
- All cluster nodes can resolve each other's hostnames correctly (check
/etc/hostsor your DNS setup) - No firewall rules are blocking default YARN ports (8032 for ResourceManager, 8042 for NodeManagers)
- Test client connectivity to the ResourceManager to rule out timeouts
One final tip: Even though I can't view your error screenshot, digging into specific error codes in YARN logs (look for ApplicationAttemptFailedEvent entries with a reason string) will help you zero in on the exact issue. For example, AM container launch failed points to node-level problems, while Resource allocation failed points to scheduler or resource shortages.
内容的提问来源于stack exchange,提问作者Shahriar Rahman

