Play Framework与Logback关联应用内存泄漏排查求助
Hey there, let's walk through how to troubleshoot this memory leak issue step by step, based on the clues you've shared:
1. Prioritize Investigating Akka Stream Timers (Your Core Clue)
The 838k+ instances of akka.stream.stage.TimerMessages$Scheduled are a huge red flag—this means thousands of uncleaned scheduled tasks are piling up in Akka Streams. Even though you aren't using timers/schedulers directly, Play relies heavily on Akka Streams under the hood (for HTTP request handling, streaming responses, WebSockets, etc.), so an internal component might be leaving timers hanging indefinitely.
- Audit Play HTTP & Akka Stream Timeout Configs: Check settings like
play.server.http.idleTimeout,play.ws.timeout.connection, or Akka'sakka.stream.materializer.timer-serviceparameters. Misconfigured timeouts or incomplete request/stream lifecycle handling can cause timers to accumulate over time. - Check for Unclosed Streams/WebSockets: If your app handles WebSockets or streaming responses, ensure these streams are properly completed/canceled when requests end. If a client disconnects but the server-side stream isn't terminated, Akka may retain timers waiting for timeouts, leading to massive buildup.
- Enable Akka Stream Debug Logs: Set the
akka.streamlogger level toDEBUGin your Logback config. This will show you which components are creating those scheduled timers (e.g., internal heartbeat tasks, timeout checks) and help pinpoint the source.
2. Dig Into the Logback THREAD_FACTORY Connection
HeapHero calling out ch.qos.logback.core.util.ExecutorServiceUtil.THREAD_FACTORY is worth exploring too. This static thread factory creates background threads for Logback (like those handling log rolling), and if those threads are stuck or holding onto object references, they can contribute to memory bloat.
- Inspect Log Rolling Behavior: Your config uses
TimeBasedRollingPolicy—could log rolling threads be blocked? Slow disk I/O, file permission issues, or stuck rollover processes can leave threads hanging, holding onto large log-related objects. Try temporarily simplifying your rolling policy or verifying disk health in your log directory. - Verify Logback Version & Compatibility: Play 2.7 ships with Logback ~1.2.3. Check for known memory leak bugs in this version; some older Logback releases had issues with unclosed appender resources. Try upgrading to the latest compatible Logback version (e.g., 1.2.12) while ensuring it works with Play 2.7's dependencies.
- Audit Trace Log Volume: You have a
performancelogger set toTRACE—if this logger generates an enormous amount of logs, your synchronous appenders might be overwhelmed, causing queue buildup and preventing object garbage collection. Consider switching to async appenders for high-volume loggers to offload the work.
3. Check for Known Play/Akka Version Bugs
Since this issue started in Play 2.6 and persists in 2.7, it's worth checking for existing bugs in the framework stack:
- Search GitHub Issues: Look through Play Framework's GitHub issues for keywords like "memory leak akka stream timer" or "TimerMessages$Scheduled". Akka 2.5.x (used in Play 2.7) had some reported timer-related leaks in specific scenarios.
- Upgrade to the Latest Play 2.7 Patch: Play 2.7.15 is the final patch release for 2.7, and it includes fixes for several memory leak issues. If you're running an older 2.7 version, upgrading to this patch might resolve the problem.
4. Enhance Heap Dump Analysis
Your current heap dump is from 10 hours of runtime, not when OOM occurs. To get clearer insights:
- Capture OOM Heap Dumps: Add JVM flags
-XX:+HeapDumpOnOutOfMemoryErrorand-XX:HeapDumpPath=/your/dump/pathto your startup script. This will generate a dump exactly when the OOM happens, showing you the exact objects consuming all memory. - Trace Reference Chains for Scheduled Timers: In tools like JXRay, VisualVM, or Eclipse MAT, pick a
TimerMessages$Scheduledinstance and trace its GC root reference chain. This will tell you exactly which component is holding onto these timers (e.g., an unclosedGraphStage, a lingering HTTP connection context). - Compare Heap Dumps Over Time: Generate dumps every 2-3 hours and compare the growth of
TimerMessages$Scheduledinstances alongside other objects. This will help you correlate the leak with specific application activity.
5. Narrow Down with Elimination Tests
- Disable Non-Core Modules: Even though you already disabled Ehcache, try turning off other non-essential Play modules (e.g., WS client, WebSocket support if unused) to see if the leak stops. This helps isolate whether the issue is in core Play or a secondary component.
- Audit Third-Party Dependencies: Check if any external libraries you're using rely on Akka Streams or Logback. A poorly implemented library might create Akka streams without properly closing them, leading to timer buildup.
内容的提问来源于stack exchange,提问作者schube

