You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

未配置Name Node High Availability时,Secondary Name Node如何处理Checkpointing操作?

Alright, let's break down exactly how the Secondary NameNode (SNN) handles checkpointing when you're running HDFS without NameNode High Availability (HA) set up. This is a critical process to keep your HDFS metadata healthy and performant, so let's walk through it step by step:

Secondary NameNode Checkpointing Mechanism (Non-HA Deployment)

Quick Background to Set the Stage

First, a quick recap of what the Active NameNode is dealing with day-to-day:

  • It keeps all HDFS metadata (file structures, permissions, block locations, etc.) in memory for fast access.
  • To persist this data, it relies on two on-disk files:
    • fsimage: A snapshot of the full metadata state from the last checkpoint. Think of this as a "full backup" of the filesystem's structure.
    • edits: A running log of every metadata change (file creates/deletes, permission updates, block allocations, etc.) that's happened since the last checkpoint.

Over time, that edits file gets huge. A massive log means longer NameNode restarts (it has to replay every single entry to rebuild the metadata) and higher risk of data loss if something goes wrong. That's where the Secondary NameNode steps in—it's like the metadata cleanup crew.

Step-by-Step Checkpointing Flow

1. When Does Checkpointing Start?

The process triggers in one of three ways:

  • Time-based: Runs automatically every X seconds, configured via dfs.namenode.checkpoint.period (default is 3600 seconds / 1 hour).
  • Size-based: Kicks off when the edits file hits a certain number of transactions, set by dfs.namenode.checkpoint.txns (default 1 million transactions).
  • Manual: You can trigger it anytime with the command hdfs dfsadmin -rollEdits (useful if you need to clean up before a restart, for example).

2. NameNode Prepares to Hand Off Work

As soon as the checkpoint is triggered, the Active NameNode does two key things to keep operations running smoothly:

  • It creates a brand new, empty edits_inprogress file. All new metadata changes from clients start going here immediately—so the NameNode never has to pause serving requests.
  • It tells the Secondary NameNode where to fetch the current fsimage and the full, unprocessed edits file from.

3. SNN Fetches and Merges the Metadata

The Secondary NameNode pulls both files to its local disk, then gets to work:

  • It loads the old fsimage into its own memory, reconstructing the metadata state as it was at the last checkpoint.
  • It replays every single transaction in the edits log on top of this in-memory state. This brings the metadata fully up to date with all changes made since the last checkpoint.

4. SNN Creates a Fresh fsimage

Once all edits are replayed, the SNN writes this merged, up-to-date metadata state to a new file on its local disk called fsimage.checkpoint. This is the new, clean snapshot of the filesystem.

5. SNN Sends the New fsimage Back to the NameNode

The Secondary NameNode uploads the fsimage.checkpoint file back to the Active NameNode.

6. NameNode Finalizes the Checkpoint

The Active NameNode wraps things up with a few file renames:

  • Renames the received fsimage.checkpoint to fsimage—replacing the old, outdated snapshot.
  • Renames the edits_inprogress file to plain old edits—since all prior changes are now safely stored in the new fsimage.
  • The old edits file gets archived or deleted (depending on your configuration) to free up space.

Key Things to Remember

  • No downtime: The NameNode never stops serving clients during checkpointing—thanks to switching to edits_inprogress right away, all new changes are captured without interruption.
  • SNN isn't a standby: In non-HA setups, the Secondary NameNode doesn't act as a failover node. Its only job is handling checkpointing to keep the NameNode's metadata files lean.
  • Risk of data loss: If the NameNode fails mid-checkpoint, any changes in edits_inprogress that haven't been merged into the new fsimage could be lost. This is one big reason HA is recommended for production clusters.

内容的提问来源于stack exchange,提问作者Rajnil Guha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:58:20