禁用向量访问边界检查时出现异常GC活动的技术问询
测试Java 19的jdk.incubator.vector和java.lang.foreign API时,发现了异常的GC行为。为避免OSR编译 artifacts,使用JMH编写了两个几乎完全相同的基准测试(示例虽刻意构造,但能复现问题):
package org.sample; import org.openjdk.jmh.annotations.*; import java.util.concurrent.TimeUnit; import java.nio.ByteOrder; import java.lang.foreign.*; import java.lang.foreign.ValueLayout; import jdk.incubator.vector.*; @OutputTimeUnit(TimeUnit.NANOSECONDS) @Warmup(iterations = 5, time = 5, timeUnit = TimeUnit.SECONDS) @Measurement(iterations = 5, time = 5, timeUnit = TimeUnit.SECONDS) @Fork(1) @State(Scope.Benchmark) public class MyBenchmark { static { // 移除该行则不会出现异常行为 System.setProperty("jdk.incubator.vector.VECTOR_ACCESS_OOB_CHECK", "0"); } final static long MEM_SIZE = Double.BYTES * 4; // SPECIES_256向量的大小 final static MemorySegment segment = MemorySegment.ofAddress(MemoryAddress.NULL, Long.MAX_VALUE, MemorySession.global()); final static ByteOrder byteOrder = ByteOrder.nativeOrder(); final long a = MemorySegment.allocateNative(MEM_SIZE, MEM_SIZE, MemorySession.global()).address().toRawLongValue(); final long b = MemorySegment.allocateNative(MEM_SIZE, MEM_SIZE, MemorySession.global()).address().toRawLongValue(); @Benchmark @BenchmarkMode(Mode.Throughput) @OperationsPerInvocation(1) public double hasNoGCActivity() { ByteVector.fromMemorySegment(ByteVector.SPECIES_256, segment, b, byteOrder) .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles() .intoMemorySegment(segment, a, byteOrder); return MemoryAddress.NULL.get(ValueLayout.JAVA_DOUBLE, a); } @Benchmark @BenchmarkMode(Mode.Throughput) @OperationsPerInvocation(1) public double hasSomeGCActivity() { ByteVector.fromMemorySegment(ByteVector.SPECIES_256, segment, b, byteOrder) .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .reinterpretAsDoubles().reinterpretAsBytes().reinterpretAsDoubles().reinterpretAsBytes() .intoMemorySegment(segment, a, byteOrder); return MemoryAddress.NULL.get(ValueLayout.JAVA_DOUBLE, a); } }
该测试需要在pom.xml中配置:
<compilerArgs> <arg>--enable-preview</arg> <arg>--add-modules</arg> <arg>jdk.incubator.vector</arg> </compilerArgs>
以及<javac.target>19</javac.target>,可通过以下命令运行:
java --add-opens java.base/jdk.internal.misc=ALL-UNNAMED --enable-preview --enable-native-access=ALL-UNNAMED --add-modules jdk.incubator.vector -verbose:gc -jar ./target/benchmarks.jar
两个基准测试几乎一致,仅第二个的链式方法调用多一次。字节重解释在ASM层面应为无操作,因此无论链式调用次数多少,二者性能应一致(交替使用Byte和Double是因为库会智能跳过同类型转换)。正常情况下二者性能确实相同,但当通过顶部的系统属性禁用{from, into}MemorySegment方法的内存访问边界检查后,第二个基准测试的性能骤降一个数量级。
-verbose:gc的输出显示,第二个方法在1-2次预热迭代后出现大量重复GC活动,这是性能下降的原因。查看生成的汇编(非OSR编译版本)发现,虽然两个测试中的冗余重解释都被优化掉,但第二个测试中存在创建Java double[]的指令,该数组地址来自本地方法外,推测是新分配的对象。
想了解这一现象的原因:是否是额外的invokevirtual字节码抑制了某些优化,导致分配无法被消除?
补充说明
查阅C2编译器的分配消除说明后,推测这可能与逃逸分析有关,额外的调用可能使分析超出了某个阈值。但在JDK19的调试版本中调整了优化器设置(如延长逃逸分析超时时间-XX:EscapeAnalysisTimeout=60),程序行为并未改变。实际上,使用-XX:+PrintEscapeAnalysis查看时,基准测试方法本身似乎从未被进行逃逸分析。
内容的提问来源于stack exchange,提问作者Francisco O.

