Selenium Grid:SO_TIMEOUT引发日志堆积致测试因内存不足终止
Let's break down your scenario first: you're running 1000+ test cases in parallel with 5 threads, and occasionally the test suite crashes due to disk exhaustion—your Selenium logs blew past 50GB! Root cause is a single session (e89da8f94eab2508559843f70803415f) spamming millisecond-level SO_TIMEOUT error logs for 7 straight hours, which filled up your Jenkins server's storage. Your core question: why didn't the Hub's timeout configurations automatically clean this rogue session up?
I'll walk through the likely causes, tied specifically to Selenium Grid 3.7.0's behavior, and share actionable fixes:
1. First, Understand Grid 3.x Timeout Logic
The Hub has two key timeout parameters that control session cleanup:
sessionTimeout: Terminates sessions that are idle (no valid test commands sent) after the set duration (default 300 seconds)browserTimeout: Forces termination of a browser instance regardless of activity, after the set max lifespan
Critical caveat: These only work reliably for sessions that can communicate normally with the Hub. If a session is in a "zombie" state (e.g., WebDriver process is unresponsive but hasn't crashed, TCP connection stays open), the Hub's detection logic might not flag it as expired.
2. Why SO_TIMEOUT Logs Spammed Your Disk
That constant SO_TIMEOUT error means the Node's WebDriver (e.g., chromedriver) couldn't communicate with the browser, or the Node lost proper heartbeat sync with the Hub. Here's what happened:
- The session entered an error loop: WebDriver kept retrying the failed connection, logging an error every millisecond
- The Hub didn't recognize this as "idle"—since logs were being generated, it might have treated the session as still active
- Over 7 hours, those tiny error lines piled up to 50GB+
3. Why Hub Timeouts Didn't Kick In
Let's dig into the specific reasons your timeout settings didn't clean up the session:
a) Timeout Configs Might Not Have Been Applied Correctly
Double-check your Hub startup command or config file. For Grid 3.7.0, you need to explicitly set the timeouts—if you relied on defaults, the 300-second sessionTimeout would never catch a 7-hour session. Example of a correct Hub startup:
java -jar selenium-server-standalone-3.7.0.jar -role hub -sessionTimeout 1800 -browserTimeout 3600
If you used a hubConfig.json, ensure both sessionTimeout and browserTimeout are set (not just one).
b) Node-Level Timeout Configs Conflicted With Hub
Grid 3.x lets Nodes define their own sessionTimeout values, which can override Hub settings. If your Node config had a longer timeout (or no explicit setting), the Hub's cleanup logic might have been ignored. Make sure Nodes inherit Hub-level timeout rules by removing any sessionTimeout lines from your Node config files.
c) The Rogue Session Didn't Trigger Idle Detection
The Hub's sessionTimeout relies on tracking "last command received" timestamps. If the session was still sending error-related traffic (even if it wasn't valid test commands), the Hub might have updated the last activity timestamp, preventing timeout cleanup. Grid 3.x doesn't distinguish between valid commands and error noise.
d) WebDriver Process Was Zombie'd
If the Node's WebDriver process (e.g., chromedriver.exe) froze but didn't crash, the Hub's termination command would never reach it. The session would stay in the Hub's registry indefinitely, and logs would keep spamming until the process was manually killed.
4. Fixes & Preventative Steps
Here's how to resolve this for your setup:
a) Lock Down Timeout Configs
- Set
browserTimeoutto a strict upper limit (e.g., 3600 seconds = 1 hour) to force-terminate any session that runs too long, regardless of activity - Set
sessionTimeoutto a reasonable idle threshold (e.g., 1800 seconds = 30 minutes) for normal test idle periods - Ensure Nodes don't override these settings—remove
sessionTimeoutfrom Node configs
b) Add Custom Session Monitoring
Grid 3.x lacks built-in zombie session detection, so add a simple script to:
- Poll the Hub's REST API (
http://<hub-url>:4444/grid/api/testsession) to list all active sessions - Check each session's creation time—if it's older than your
browserTimeout, send a DELETE request to terminate it
c) Limit Log Growth
Add log rotation to your Jenkins setup to cap Selenium log file sizes (e.g., 1GB per file) and delete old logs automatically. This prevents a single rogue session from taking down your server.
d) Fix the Root SO_TIMEOUT Issue
- Match your WebDriver version to your browser version (Grid 3.7.0 works best with ChromeDriver 2.33–2.38 for Chrome 60–65, adjust for your browser)
- Check test cases for unoptimized waits (avoid
Thread.sleep()—use explicit waits instead) that might leave sessions idle too long - Ensure Node servers have enough resources (CPU/RAM) to avoid WebDriver-browser communication delays
5. Your Provided Error Log & Configs
Error Log Snippet
> [e89da8f94eab2508559843f70803415f] ERROR org.openqa.selenium.remote.server.DriverServlet - Exception executing websocket request > java.net.SocketTimeoutException: SO_TIMEOUT expired > at org.openqa.selenium.remote.server.netty.NettyWebSocketHandler.channelRead0(NettyWebSocketHandler.java:97) > at io.netty.channel.SimpleChannelInboundHandler.channelRead(SimpleChannelInboundHandler.java:105) > ... (truncated stack trace)
Hub Configuration (hubConfig.json)
{ "port": 4444, "newSessionWaitTimeout": -1, "cleanUpCycle": 5000, "role": "hub", "browserTimeout": 3600, "sessionTimeout": 1800 }
Node Configuration (nodeConfig.json)
{ "capabilities": [ { "browserName": "chrome", "maxInstances": 5, "platform": "WINDOWS", "version": "114.0" } ], "port": 5555, "registerCycle": 5000, "hub": "http://hub-ip:4444", "role": "node", "cleanUpCycle": 5000, "timeout": 300, "sessionTimeout": 1800 }
内容的提问来源于stack exchange,提问作者Vignesh Paramasivam

