关于CMS GC初始标记阶段年轻代扫描方式的技术问询
Great question—this is a nuanced detail that trips up many folks learning CMS! Let’s break down exactly how CMS handles young-to-old generation references during the initial mark phase, and how it differs from a full Minor GC scan.
First, let’s recap the goal here: during initial mark (the first short STW pause in CMS), we need to mark two sets of old-gen objects:
- Objects directly reachable from GC Roots
- Old-gen objects that are referenced by still-alive young-gen objects (critical because CMS collects the old gen independently)
The key thing to understand is that CMS does NOT perform a full scan of the young generation here—that would defeat the purpose of CMS’s low-latency design, since a full young-gen scan would lengthen the STW pause. Instead, it relies on a helper data structure called the Card Table.
Here’s the step-by-step breakdown:
- The Card Table divides the old generation into small, fixed-size "cards" (usually 512 bytes each).
- Whenever a young-gen object references an old-gen object, the JVM uses a write barrier to mark the corresponding card in the table as "dirty". This happens in real time as your application runs, not just during GC.
- During the initial mark STW pause:
- First, CMS quickly marks all old-gen objects directly reachable from GC Roots.
- Then, instead of scanning the entire young gen, it iterates only over the dirty cards in the Card Table. For each dirty card, it identifies the old-gen objects that are referenced by young-gen objects and marks those as alive.
This approach keeps the initial mark pause extremely short because it only processes a tiny subset of the old gen (the dirty cards) instead of the entire young generation.
To contrast with a Minor GC: a Minor GC does scan the entire young generation (Eden + Survivor spaces) to mark all alive objects, then copies those survivors to a new space. That’s a full, time-consuming scan—something CMS avoids in its initial mark phase to preserve low latency.
It’s worth noting that this reliance on the Card Table is why CMS has some overhead during application runtime (due to the write barrier checks), but that’s the tradeoff for its low-pause garbage collection in the old gen.
内容的提问来源于stack exchange,提问作者Noskol

