You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

响应超时测试中Jetty线程阻塞致应用崩溃问题排查

Root Cause Analysis

The core issue is a bottleneck in uncaught exception handling that triggers thread blocking, resource exhaustion, and eventual OOM. Here's the breakdown:

  1. Massive Uncaught Exception Volume: With 1000 TPS of timed-out requests, every failure generates an exception that propagates to Reactor's scheduler. Since no custom exception handler is configured, the JVM's default ThreadGroup.uncaughtException handler takes over.
  2. PrintStream Lock Contention: The default handler writes stack traces to System.err (a synchronized PrintStream). One thread (#659) holds this lock while stuck in a slow native method (StackTraceElement.initStackTraceElements) generating a stack trace. Hundreds of other threads (#660 and more) block waiting to acquire the same lock, creating a critical bottleneck.
  3. Resource Exhaustion: Blocked threads consume memory (stack space, thread metadata). Jetty's thread pool expands to handle the backlog, worsening memory pressure until the JVM hits OOM.
Why This Occurs in Your Setup
  • Unbounded Exception Flow: 1000 exceptions per second funnel through a single synchronized resource, creating a queue of blocked threads.
  • Slow Stack Trace Generation: The native stack trace generation process is computationally expensive, and holding the PrintStream lock during this operation amplifies the bottleneck.
  • Missing Custom Error Handling: Reactor/WebFlux relies on the JVM's default handler instead of a non-blocking, async-aware alternative.
Solutions

1. Implement a Custom Uncaught Exception Handler

Replace the default handler to avoid synchronized PrintStream usage. Configure Reactor's scheduler with a non-blocking logger:

Schedulers.onHandleError((thread, throwable) -> {
    LoggerFactory.getLogger(thread.getClass())
        .error("Uncaught exception in thread: {}", thread.getName(), throwable);
});

Use an async logging appender (e.g., Logback AsyncAppender) to prevent logging from blocking threads.

2. Cap Jetty Thread Pool Size

Prevent unbounded thread creation by setting explicit limits on Jetty's thread pool:

HttpClient httpClient = new HttpClient(new HTTP2Client());
httpClient.setExecutor(new QueuedThreadPool(200, 50)); // Adjust based on your workload

This limits memory consumption even during contention.

3. Add Graceful Timeout Handling in WebFlux

Catch timeouts before they become uncaught exceptions to reduce the volume of errors reaching the scheduler:

webClient.get()
    .uri("/endpoint")
    .retrieve()
    .bodyToMono(YourResponse.class)
    .timeout(Duration.ofSeconds(5))
    .onErrorResume(TimeoutException.class, ex -> {
        log.warn("Request timed out", ex);
        return Mono.just(fallbackResponse); // Return a fallback or empty response
    });

4. Optimize Stack Trace Generation (Optional)

If stack traces are necessary for debugging:

  • Disable stack traces for known, non-critical exceptions (like timeouts) to reduce processing overhead.
  • Use a logger that caches or processes stack traces asynchronously.
Verification

After applying these changes:

  • Thread blocking on PrintStream will no longer occur.
  • The application will handle 1000 TPS of timeouts without crashing.
  • Memory usage will remain stable under failure conditions.

内容的提问来源于stack exchange,提问作者owais nazir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 09:54:49