Hadoop 2.7.5+HBase五节点分布式集群:监控告警与开机自启咨询
Hey there! Let's break down your two questions step by step—monitoring daemon failures with alerts, and setting up Hadoop/HBase to start on boot. I've worked with Hadoop 2.7.5 clusters a lot, so here's what I recommend:
First off, neither Hadoop 2.7.5 nor HBase have a fully built-in "alert me when a daemon dies" system out of the box, but they give you all the tools you need to build one yourself, plus there's a semi-official tool that makes it easier.
1. Build Alerts Using Hadoop's Native Monitoring Interfaces
All Hadoop daemons (NameNode, ResourceManager, DataNode, etc.) and HBase components (HMaster, RegionServer) expose JMX metrics endpoints with default ports (e.g., NameNode on 50070, HMaster on 16010). You can write simple scripts to poll these endpoints and trigger emails if something looks off.
Here's a quick shell script example to check if your NameNode is alive—you can run this via cron every 5 minutes:
# Check if NameNode's HTTP endpoint is reachable and returns valid status curl -s http://your-namenode-host:50070/jmx | grep -q "NameNodeStatus" if [ $? -ne 0 ]; then # Send email alert (you'll need mailutils or a similar service configured) echo "URGENT: NameNode on your-namenode-host is unresponsive!" | mail -s "Hadoop Alert: NameNode Down" your-alert-email@example.com fi
2. Use HBase's Built-in Health Check Tool
HBase has the hbase hbck command, which scans the entire cluster for issues like missing RegionServers or unresponsive HMasters. You can wrap this in a script to trigger alerts on errors:
# Run HBase health check HBCK_RESULT=$(hbase hbck) if echo "$HBCK_RESULT" | grep -q "ERROR"; then echo "HBase Cluster Issues Detected: $HBCK_RESULT" | mail -s "HBase Alert: Cluster Health Error" your-alert-email@example.com fi
3. Semi-Official Tool: Ambari (Easier Alternative)
If you don't want to maintain custom scripts, Ambari is the way to go. It's a Hadoop cluster management tool that supports Hadoop 2.7.5, and it has built-in monitoring, alerting, and even auto-recovery features. You can configure email alerts directly in the UI for when any daemon goes down—no scripting required. It's not "built into" Hadoop itself, but it's the de facto standard tool for this kind of management in the ecosystem.
The most reliable way to do this is using systemd (for CentOS 7+/Ubuntu 16.04+ systems)—it's the modern init system and handles process restarting if needed too. Here's how to set it up:
1. Create Systemd Service Files for Hadoop Daemons
Let's start with the NameNode. Create a file at /etc/systemd/system/hadoop-namenode.service with this content:
[Unit] Description=Hadoop NameNode After=network.target [Service] User=hadoop # Replace with your Hadoop runtime user Group=hadoop ExecStart=/path/to/your/hadoop/bin/hdfs --daemon start namenode ExecStop=/path/to/your/hadoop/bin/hdfs --daemon stop namenode Restart=on-failure # Optional: auto-restart if the daemon crashes [Install] WantedBy=multi-user.target
Repeat this for other Hadoop daemons—just tweak the Description, ExecStart, and ExecStop lines:
- DataNode:
hdfs --daemon start datanode - ResourceManager:
yarn --daemon start resourcemanager - SecondaryNameNode:
hdfs --daemon start secondarynamenode
2. Create Systemd Service Files for HBase
For the HMaster, create /etc/systemd/system/hbase-hmaster.service:
[Unit] Description=HBase HMaster After=network.target hadoop-namenode.service hadoop-resourcemanager.service [Service] User=hadoop Group=hadoop ExecStart=/path/to/your/hbase/bin/hbase-daemon.sh start master ExecStop=/path/to/your/hbase/bin/hbase-daemon.sh stop master Restart=on-failure [Install] WantedBy=multi-user.target
For RegionServers, create a similar file named hbase-regionserver.service, changing the command to hbase-daemon.sh start regionserver and adding hbase-hmaster.service to the After line.
3. Enable & Start the Services
Once your service files are ready, run these commands to activate them:
# Reload systemd to pick up the new service files sudo systemctl daemon-reload # Enable the services to start on boot sudo systemctl enable hadoop-namenode.service sudo systemctl enable hadoop-datanode.service sudo systemctl enable hbase-hmaster.service # Add all other daemons you created here # Start the services immediately (so you don't have to reboot) sudo systemctl start hadoop-namenode.service sudo systemctl start hadoop-datanode.service
Quick Note for Older Systems
If you're running an older OS (like CentOS 6) that uses init.d instead of systemd, you can write simple init scripts that call the Hadoop/HBase daemon start commands, place them in /etc/init.d/, and use chkconfig to enable them on boot.
内容的提问来源于stack exchange,提问作者Hafiz Muhammad Shafiq

