使用Selenium 4 Docker Grid镜像并行运行Gauge时随机触发TimeoutException问题求助
Hey there, let’s work through this random flakiness you’re hitting when running Gauge tests in parallel with Selenium 4 Docker Grid (versions selenium/node-chrome:4.3.0-20220706 and selenium/hub:4.3.0-20220706). The stack trace points to Netty HTTP handler failures, which often tie to resource constraints, timeouts, or version-specific bugs. Here are actionable fixes to try:
Fix Docker Node Resource Constraints
Chrome is resource-heavy, and the default Selenium Docker nodes don’t allocate enough shared memory or CPU/RAM for parallel runs. This causes random crashes or timeouts.- Add
--shm-size=2gto your node run command (Chrome relies on shared memory to avoid crashes):docker run -d --shm-size=2g --link selenium-hub:hub selenium/node-chrome:4.3.0-20220706 - Limit CPU/RAM per node to prevent resource starvation, e.g.,
--cpus=2 --memory=4g - Check the Grid UI (
http://<hub-ip>:4444/ui) for node health, and inspect node logs for Chrome crash messages (docker logs <node-container-id>)
- Add
Tune Selenium Client Timeouts
The Netty errors often stem from unresponsive requests. Adjust your RemoteWebDriver timeout settings to account for parallel grid delays:ChromeOptions options = new ChromeOptions(); // Add your desired capabilities // Configure custom HTTP client with longer timeouts HttpClient.Factory httpClientFactory = HttpClient.Factory.create() .withConnectTimeout(Duration.ofSeconds(15)) .withReadTimeout(Duration.ofSeconds(15)); HttpCommandExecutor executor = new HttpCommandExecutor(URI.create("http://hub:4444/wd/hub"), httpClientFactory); RemoteWebDriver driver = new RemoteWebDriver(executor, options); // Set standard driver timeouts driver.manage().timeouts().implicitlyWait(Duration.ofSeconds(10)); driver.manage().timeouts().pageLoadTimeout(Duration.ofSeconds(30)); driver.manage().timeouts().scriptTimeout(Duration.ofSeconds(10));Optimize Gauge Parallel Execution
Don’t overwhelm your Grid with more concurrent tests than it can handle:- In
env/default/default.properties, setmax_concurrencyto match your Grid’s capacity (e.g., if you have 2 nodes each supporting 4 sessions, set it to 8):parallel=true max_concurrency=8 - Ensure every test scenario cleans up its WebDriver instance in an
AfterScenariohook, even if the test fails. Leaking drivers will clog the Grid and cause random failures.
- In
Upgrade Selenium Docker Images
The 4.3.0-20220706 version is quite old and has known stability issues with parallel runs. Upgrade to a recent stable release (e.g.,4.15.0-20231108) to pick up bug fixes for Netty and Grid session management:docker run -d -p 4444:4444 selenium/hub:4.15.0-20231108 docker run -d --shm-size=2g --link selenium-hub:hub selenium/node-chrome:4.15.0-20231108Add Retry Logic for Flaky Operations
Since failures are random, wrap high-risk operations (clicks, element finds, switches) in retry logic. You can use Gauge’s built-in retry annotation or custom code:- Gauge annotation example:
@Retry(count = 2) fun selectAffiliateFromReport() { // Your step logic here } - Custom retry wrapper for element actions:
fun checkOrUncheckCheckbox(element: RemoteWebElement, shouldCheck: Boolean) { val maxAttempts = 3 var attempt = 0 while (attempt < maxAttempts) { try { if (element.isSelected != shouldCheck) { element.click() } break } catch (e: Exception) { attempt++ if (attempt == maxAttempts) throw e Thread.sleep(1000) // Short wait before retrying } } }
- Gauge annotation example:
Rule Out Network Issues
Random Netty failures can also come from unstable network connections between your test runner and Grid:- Test connectivity with
ping <hub-ip>ortraceroute <hub-ip>to check for latency or packet loss - If using Docker bridge network, try host network mode (if your environment allows) to eliminate port mapping overhead:
docker run -d --network host selenium/hub:4.3.0-20220706 docker run -d --network host --shm-size=2g selenium/node-chrome:4.3.0-20220706
- Test connectivity with
内容的提问来源于stack exchange,提问作者mbdev

