NameNode如何识别故障DataNode的数据块及副本维持处理流程?
Great question—this cuts to how HDFS keeps data reliable without making the NameNode grind to a halt. Let’s break this down clearly:
How NameNode Quickly Identifies All Block Copies on a Failed DataNode (No FSImage Traversal Needed)
The key here is that the NameNode doesn’t rely on the FSImage (a persistent on-disk snapshot) for real-time node-block tracking. Instead, it maintains in-memory data structures that let it instantly pull up the blocks on any dead node:
- Every DataNode sends a full Block Report to the NameNode when it first boots up, listing all blocks stored on that node.
- After the initial report, the DataNode sends incremental updates (Incremental Block Reports) whenever blocks are added or removed.
- The NameNode keeps a dedicated mapping: for each active DataNode, it stores an in-memory set of all blocks that node holds. This is part of the NameNode’s core runtime state, not something it has to read from disk.
When a DataNode is marked dead (after missing heartbeats beyond the configured timeout, default ~10 minutes), the NameNode just looks up this pre-existing in-memory set for that node. No need to scan the entire FSImage—this is an almost instant lookup to get all blocks on the failed node.
Step-by-Step Recovery Process
Here’s the play-by-play once a node is flagged as dead:
- Mark the node as dead: The NameNode’s heartbeat monitor detects missing heartbeats and moves the DataNode from the "active" pool to the "dead" pool.
- Fetch the node’s block list: Pull the pre-stored set of blocks from the in-memory DataNode-to-block map.
- Update block replica counts: For each block in that list, decrement its current replica count in the NameNode’s block metadata.
- Check for under-replicated blocks: Compare the updated replica count against the configured replication factor (default 3). Any block with fewer replicas than required is marked as "under-replicated".
- Schedule replica creation: The NameNode’s replication scheduler picks these under-replicated blocks and assigns tasks to surviving DataNodes to copy the blocks from existing live replicas, restoring the replica count to the desired level.
How NameNode Tracks Which Blocks Are on a Specific DataNode
This ties back to the initial and incremental block reports:
- When a DataNode starts, it scans its local storage for all blocks and sends the full list to the NameNode. The NameNode adds each block to the node’s in-memory set.
- When a block is written to the DataNode (as part of a new file or replication), the DataNode sends an incremental report to the NameNode, which adds the block to the node’s set.
- When a block is deleted (e.g., file deletion, over-replication cleanup), the DataNode sends another incremental report, and the NameNode removes the block from the node’s set.
At any time, the NameNode can instantly answer "what blocks are on DataNode X?" by checking this in-memory set—no disk I/O or full metadata scans required.
内容的提问来源于stack exchange,提问作者arctic_Oak

