对象创建是否影响应用性能?高创建率下GC正常的性能困惑
Hey there, let's unpack your problem—this is a super common gotcha when debugging performance, because "good GC metrics" don't always tell the whole story. Let's break down why your high object creation rate is still tanking performance under load, especially with that wild discrepancy between single-user (2s) and 100-user (2min) tests:
1. Object Creation Has Hidden Costs Beyond GC
Even if your GC logs look clean after 20 minutes, the act of creating objects itself carries overhead that adds up fast under concurrency:
- Allocation Contention: In most JVMs, young-gen (Eden) allocation uses a pointer updated with CAS operations. With 100 concurrent threads all trying to allocate objects, this creates constant CAS contention—something you’ll never see in a single-user test where there’s no competition for that pointer.
- TLB & Cache Pressure: Every object allocation requires memory address translation. When hundreds of threads are allocating at once, you’ll hit frequent TLB (Translation Lookaside Buffer) misses, forcing the CPU to fetch data from main memory instead of fast cache. This alone can slow things down drastically.
- Object Initialization Overhead: Zeroing out object fields, setting up object headers (mark word, class pointer)—these are tiny per-object, but multiply by millions of objects under load, and they become a significant CPU drain.
2. Escape Analysis Might Be Failing Under Concurrency
You mentioned these code segments are in-memory only (no DB calls), which should make them prime candidates for escape analysis optimizations like stack allocation or scalar replacement. But here’s the catch:
- Escape analysis works best in single-threaded or simple concurrent scenarios. If your code has subtle shared references, branching that the JVM can’t analyze, or even just complex control flow, the JVM might bail out and allocate objects on the heap instead.
- In your single-user test, escape analysis likely kicks in perfectly—objects are allocated on the stack (no heap overhead, no GC), hence the 2s runtime. But under 100 users, the JVM can’t optimize as effectively, so you’re stuck with heap allocations and all the contention that comes with them.
3. "Good GC" Doesn’t Mean No Hidden Pauses
Your GC report says performance is "not bad" after 20 minutes, but GC has subtle overhead that doesn’t show up as full stop-the-world pauses:
- Minor GC Overhead: Even fast Minor GCs have overhead for scanning young-gen, updating references, and coordinating between threads. Under heavy allocation, these GCs fire more frequently, and the cumulative delay adds up across 100 concurrent requests.
- Allocation Stalls: If Eden fills up faster than the GC can clean it, threads might stall waiting for space to be freed. These stalls are often too short to show up as explicit GC pauses in logs, but they kill throughput under load.
Next Steps to Validate Your Suspicions
To confirm this is the root cause, try these actionable steps:
- Use a profiler like AsyncProfiler to check for CAS contention or memory allocation hotspots. Look for high CPU time in
mallocor JVM allocation routines. - Enable escape analysis logging with
-XX:+PrintEscapeAnalysisto see if your target code is actually getting optimized. If you see lots of "escaped" objects that you didn’t expect, that’s a red flag. - Test object pooling for the most frequently created objects. Even a simple pool can eliminate allocation contention and cut down on memory overhead.
- Compare the number of objects created per request in single-user vs 100-user tests. You might find that concurrency is causing more objects to be created per request than you think (e.g., thread-local objects that aren’t being reused).
内容的提问来源于stack exchange,提问作者Mkp

