将代码移至平台线程运行以避免虚拟线程固定
问题描述
我正在将Spring Boot应用从平台线程迁移至虚拟线程,整体运行正常,但部分请求中第三方库的同步块内阻塞操作会导致虚拟线程固定,若不解决可能出现所有载体线程被固定、无可用平台线程处理请求的情况。
在等待第三方库修复期间,我计划采用如下方案:将导致虚拟线程载体固定的代码提交至基于平台线程的ExecutorService,再从虚拟线程中立即调用返回的CompletableFuture的.get()方法。
以下是演示代码(Spring Boot):
import org.springframework.boot.ApplicationArguments; import org.springframework.boot.ApplicationRunner; import org.springframework.stereotype.Component; import java.time.Instant; import java.util.concurrent.ExecutorService; import java.util.concurrent.Executors; @Component class AppRunner implements ApplicationRunner { private final Object monitor = new Object(); private final ExecutorService executorService; public AppRunner() { this.executorService = Executors.newFixedThreadPool(2); } @Override public void run(ApplicationArguments args) { String parallelismNum = System.getProperty("jdk.virtualThreadScheduler.parallelism"); System.out.println("jdk.virtualThreadScheduler.parallelism is set to: " + parallelismNum); try (var virtualExecutor = Executors.newVirtualThreadPerTaskExecutor()) { for (int taskId = 0; taskId < 100; taskId++) { final int taskIdentifier = taskId; System.out.println("Submitting the task to executor: " + taskIdentifier); virtualExecutor.submit(() -> simpleLongOperation(taskIdentifier)); /** * If I submit the following instead of pinningOperationOnPlatform, * My simpleLongOperation would be waiting for a carrier to unpin */ // virtualExecutor.submit(() -> pinningOperation(taskIdentifier)); virtualExecutor.submit(() -> pinningOperationOnPlatform(taskIdentifier)); } } } /** * This runs the pinning operation on a platform thread but awaits on the virtual. * So no pinning is actually happening. */ private void pinningOperationOnPlatform(int id) { try { executorService.submit(() -> pinningOperation(id)) .get(); } catch (Exception e) { throw new RuntimeException(e); } } private void simpleLongOperation(int id) { System.out.println("\tSIMPLE Start operation " + id + " : " + Instant.now()); sleep(1_000); System.out.println("\tSIMPLE Complete operation " + id + " : " + Instant.now()); } private void pinningOperation(int id) { System.out.println("PINNING Start operation " + id + " : " + Instant.now()); synchronized (monitor) { sleep(5_000); } System.out.println("PINNING Complete pinning operation " + id + " : " + Instant.now()); } private void sleep(long millis) { try { Thread.sleep(1_000); } catch (InterruptedException e) { throw new RuntimeException(e); } } }
我的理解是:该操作会在执行器池线程上等待进入同步块,但由于.get()从虚拟线程调用,操作完成后载体线程会被释放。目前该方案仅对需要执行该操作的少量请求限制并行度为线程池大小,我认为对应用性能影响不大,请问我的假设是否正确,有无遗漏点?
分析与结论
你的核心假设是正确的,这个临时方案确实能有效避免虚拟线程固定问题,且对整体性能影响可控,但有几个需要注意的遗漏点:
1. 线程池参数配置
- 使用
Executors.newFixedThreadPool(2)的固定大小线程池,在请求量突增时会导致大量调用.get()的虚拟线程进入等待状态。虽然虚拟线程本身开销极低,但过多等待的虚拟线程仍会占用少量内存(每个虚拟线程栈初始仅几KB),如果这类请求量级极大,需评估内存占用情况。 - 建议根据这类请求的实际并发量调整线程池大小,或使用
ThreadPoolExecutor自定义拒绝策略(比如队列满时抛出异常或降级处理),避免无限制的任务堆积。
2. 异常处理与中断传递
- 当前代码中
get()抛出的异常被包装为RuntimeException,但虚拟线程的中断信号不会自动传递给平台线程池中的任务。如果虚拟线程被中断(比如请求超时),平台线程池中的pinningOperation仍会继续执行,造成资源浪费。 - 可通过
Future.cancel(true)传递中断信号,优化后的代码示例:
private void pinningOperationOnPlatform(int id) { var future = executorService.submit(() -> pinningOperation(id)); try { future.get(); } catch (InterruptedException e) { Thread.currentThread().interrupt(); // 保留中断状态 future.cancel(true); // 中断平台线程上的任务 throw new RuntimeException("Operation interrupted", e); } catch (ExecutionException e) { throw new RuntimeException("Operation failed", e.getCause()); } }
3. 资源释放
- 当前代码中
executorService是类成员变量,未被正确关闭。在Spring Boot应用中,应将其声明为@Bean并指定destroyMethod="shutdown",或在@PreDestroy方法中关闭线程池,避免应用关闭时出现资源泄漏:
@PreDestroy public void shutdownExecutor() { executorService.shutdown(); try { if (!executorService.awaitTermination(60, TimeUnit.SECONDS)) { executorService.shutdownNow(); } } catch (InterruptedException e) { executorService.shutdownNow(); Thread.currentThread().interrupt(); } }
4. 性能权衡边界
- 少量请求场景下,该方案的性能损耗可忽略——虚拟线程等待
.get()时会释放载体线程,不会占用平台线程资源。但如果这类请求占比提升,平台线程池的并行度会成为瓶颈,此时需重新评估线程池大小,或考虑其他临时方案(若无法修改第三方库代码,线程池方案仍是最优选择)。
5. 虚拟线程调度的隐性影响
- 大量虚拟线程同时等待平台线程池任务完成,可能导致虚拟线程调度器的任务队列变长,但由于虚拟线程调度开销极低,这种情况一般不会造成明显性能问题,仅在极端高并发场景下需要关注。
总的来说,这个临时方案合理且有效,只要注意上述细节,就能在第三方库修复前稳定运行。
内容的提问来源于stack exchange,提问作者Andrey Antipov
相关产品推荐
相关产品推荐

