Oracle HotSpot客户端与服务端模式性能对比实践探究
Awesome that you’re diving deep into HotSpot’s inlining and branch prediction—these are core pieces of what makes the JVM so performant, and testing the Client vs Server mode differences is a great way to cement your understanding. Let’s walk through how to build and analyze those tests step by step.
First, Align Your Test Goals
Since you already have a foundational grasp of inlining and branch prediction, your tests should target how these mechanisms behave differently across the two modes:
- Client Mode (C1 Compiler): Prioritizes fast startup with shallow inlining thresholds and basic branch prediction optimizations.
- Server Mode (C2 Compiler): Optimizes for long-running performance with aggressive, deep inlining (including cross-level method calls) and branch prediction that leverages runtime execution statistics.
Build Targeted Test Programs
You’ll want two distinct test cases to isolate the impact of inlining and branch prediction separately.
1. Inlining Performance Test
This test uses nested method calls to highlight how aggressively each compiler eliminates method-invocation overhead via inlining. C2 will far outperform C1 here by collapsing the nested calls into direct computation.
public class InliningTest { private static int deepNestedCall(int depth) { if (depth == 0) return 1; return deepNestedCall(depth - 1) * 2; } public static void main(String[] args) { // Warm up the JVM to trigger JIT compilation for (int i = 0; i < 100_000; i++) { deepNestedCall(10); } // Run the timed test long startTime = System.nanoTime(); for (int i = 0; i < 1_000_000; i++) { deepNestedCall(10); } long endTime = System.nanoTime(); System.out.printf("Total execution time: %d ms%n", (endTime - startTime) / 1_000_000); } }
Run it with both modes (note: Java 8+ defaults to Server mode on most platforms, so explicit flags are needed):
- Client Mode:
java -client InliningTest - Server Mode:
java -server InliningTest
2. Branch Prediction Test
This test compares performance between predictable and unpredictable branches. C2 uses runtime data to optimize predictable branches far more effectively than C1, so you’ll see a bigger performance gap between ordered and unordered data in Server mode.
import java.util.Random; public class BranchPredictionTest { public static void main(String[] args) { int[] array = new int[1_000_000]; Random random = new Random(42); // Fixed seed for consistency // Create an ordered array (predictable branches) for (int i = 0; i < array.length / 2; i++) { array[i] = random.nextInt(100); } for (int i = array.length / 2; i < array.length; i++) { array[i] = random.nextInt(100) + 1000; } // Test predictable branch long startTime = System.nanoTime(); int sum = 0; for (int num : array) { if (num > 500) { sum += num; } } long endTime = System.nanoTime(); System.out.printf("Ordered branch time: %d ms | Sum: %d%n", (endTime - startTime) / 1_000_000, sum); // Shuffle array for unpredictable branches shuffleArray(array); // Test unpredictable branch startTime = System.nanoTime(); sum = 0; for (int num : array) { if (num > 500) { sum += num; } } endTime = System.nanoTime(); System.out.printf("Unordered branch time: %d ms | Sum: %d%n", (endTime - startTime) / 1_000_000, sum); } private static void shuffleArray(int[] array) { Random random = new Random(42); for (int i = array.length - 1; i > 0; i--) { int index = random.nextInt(i + 1); int temp = array[index]; array[index] = array[i]; array[i] = temp; } } }
Run this with the same -client and -server flags, and focus on the gap between ordered and unordered execution times.
Analyze Your Results Effectively
- Inlining Test: Server mode should have significantly lower execution time—this is C2 eliminating the nested method call overhead via deep inlining. You can verify this by adding JIT logging:
java -server -XX:+PrintCompilation -XX:+PrintInlining InliningTestto see exactly which methods get inlined. - Branch Prediction Test: Look at the difference between ordered and unordered times. Server mode will show a much larger gap, as C2 uses runtime branch statistics to predict predictable branches almost perfectly, while C1’s optimization is more basic.
- Repeat Tests: Run each test multiple times and average the results to account for JVM warm-up and system-level variability.
Quick Notes for Modern Java Versions
If you’re using Java 10+, the Client mode has been removed. To simulate C1-only compilation (similar to legacy Client mode), use the flag: -XX:TieredStopAtLevel=1.
内容的提问来源于stack exchange,提问作者Paul Warnick

