You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

G1 GC新生代Ext Root Scanning耗时过高如何排查?

Hey there, let's tackle that G1 GC Ext Root Scanning bottleneck you're seeing—those numbers (maxing out at 137ms!) are definitely a red flag, especially since they're eating into your app's responsiveness. I've debugged similar issues before, so here are some practical, actionable steps to get to the bottom of it:

1. Validate Your Code Cache Hypothesis First

You mentioned your code cache is around 90MB, and that's a great starting point because G1 scans the code cache as part of root scanning. A bloated code cache means more entries to check, so let's test this first:

  • Try limiting the code cache size with JVM flags to see if it moves the needle:
    -XX:InitialCodeCacheSize=64m -XX:ReservedCodeCacheSize=64m
    
    Start with a lower value than your current 90MB (adjust based on your app's actual needs—you don't want to hit code cache full errors). Monitor if Ext Root Scanning times drop after this change; if they do, your code cache was holding too much unused compiled code.
  • Enable code cache logging to see what's filling it up:
    • Use -XX:+PrintCodeCache to log usage details when the JVM exits.
    • Use -XX:+PrintCodeCacheOnCompilation to track individual compilations and see which methods are taking up space.
2. Break Down Root Scanning Time with Detailed GC Logs

Ext Root Scanning covers multiple root types—don't just fixate on code cache. Let's get granular data on where the time is going:

  • Enable detailed G1 GC logging with these flags:
    -XX:+PrintGCDetails -XX:+PrintGCTimeStamps -XX:+PrintGCApplicationStoppedTime -XX:+PrintReferenceGC
    
  • Look for sections in the log that break down root scanning into categories like Code Roots, Classloader Roots, JNI Roots, or Stack Roots. This will tell you exactly which root type is eating up the most time.
  • For example:
    • If Stack Roots are high, you might have too many threads (each thread's stack gets scanned) or extremely deep call stacks. Use jstack <pid> to check thread counts—idle threads add unnecessary scanning overhead.
    • If Classloader Roots are high, that points to classloader bloat or leaks.
3. Check for Classloader Leaks or Bloat

Classloaders and their loaded classes are part of the root set, so leaks here can drastically increase root scanning time:

  • Use jmap -histo:live <pid> to count instances of java.lang.Class and java.lang.ClassLoader. If these numbers keep growing over time, you've got a leak (common in apps with hot deployment like web apps).
  • Take a heap dump with jmap -dump:live,format=b,file=heap.hprof <pid> and analyze it with tools like VisualVM or Eclipse MAT. Look for classloaders that are retained in memory but shouldn't be (e.g., old web app classloaders after redeployment).
4. Audit JNI Usage

JNI roots (native references to Java objects) are another often-overlooked culprit. If your app uses a lot of JNI calls, especially with global references, these add to the root scan workload:

  • Enable JNI logging with -Xcheck:jni—this logs all JNI calls and flags issues like incorrect global reference management.
  • Check if global references are being leaked (not deleted when no longer needed). Every unused global reference is an extra root G1 has to scan.
5. Use Profilers for Deep Dives

Sometimes GC logs aren't enough—profilers can show you exactly where the JVM is spending time during root scanning:

  • AsyncProfiler is a lightweight, low-overhead option. Run it to capture root scanning events with:
    ./profiler.sh -d 30 -e gc-root-scanning <pid>
    
    This will generate a flame graph showing hot spots in the root scanning process.
  • Tools like JProfiler or YourKit can also attach to the JVM and trace GC pause details, helping you spot unexpected root sources (like a huge number of weak references that need processing).
6. Compare Across Your Nodes

You mentioned multiple nodes running the same app—use this to your advantage:

  • Check if Ext Root Scanning times are consistent across all nodes. If only some nodes are worse, look for environmental differences: different JVM versions, varying load levels, OS settings, or even different deployment configurations.
  • Test changes (like adjusting code cache size) on one node first before rolling out to your entire fleet. This lets you validate fixes without risking downtime for all users.

内容的提问来源于stack exchange,提问作者Laxman Prabhu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:12:07