Selenium HtmlUnit WebDriver报错10061:连接被主动拒绝
Hey there! Let's break down why your switch from Chrome WebDriver to HtmlUnit Remote WebDriver is causing problems, even after granting firewall access to python.exe. Here are the key steps you might have missed:
1. Verify HtmlUnit Server is Running & Version Compatible
- Unlike Chrome Driver (which launches a browser process directly), HtmlUnit Remote requires you to start the HtmlUnit Selenium Server first on port 3000. If you haven't done this, your Python code can't connect to anything at
http://127.0.0.1:3000/wd/hub. - Double-check version matching: Ensure your Selenium Python binding version aligns with the HtmlUnit Server version. Mismatched versions (e.g., Selenium 4.x with an old HtmlUnit Server) often cause connection failures or unrecognized commands.
2. Adjust Desired Capabilities for Better Compatibility
Your current capability uses plain HTMLUNIT, which simulates an older, generic browser. This can trigger errors or blocks on modern sites like Google. Try updating your code to use a browser-specific capability:
from selenium import webdriver # Use Chrome-emulating capabilities for better site compatibility caps = webdriver.DesiredCapabilities.HTMLUNIT_WITH_CHROME.copy() caps["javascriptEnabled"] = True # Explicitly enable JS (critical for most sites) caps["version"] = "latest" # Simulate the latest Chrome version driver = webdriver.Remote("http://127.0.0.1:3000/wd/hub", caps) driver.get("https://google.com/")
3. Check Firewall Access for the HtmlUnit Server Process
You granted access to python.exe, but the HtmlUnit Server is a Java process (usually java.exe). Make sure your firewall allows this process to send/receive network traffic, especially on port 3000. A quick test: temporarily disable your firewall and see if the connection works. If it does, add a firewall rule for the Java process.
4. Confirm Port 3000 is Being Listened To
Open Command Prompt and run this command to check if the HtmlUnit Server is actually using port 3000:
netstat -ano | findstr :3000
If no results appear, the server isn't running correctly. Verify your startup command (example for latest HtmlUnit Server):
# Replace with your actual jar filename java -jar htmlunit-driver-standalone-4.9.0.jar -port 3000
Check the server's console output for errors like port conflicts or missing Java dependencies (HtmlUnit requires Java 11+).
5. Test with a Simpler Website First
Google has strict anti-scraping measures that might block HtmlUnit's default configuration. Try accessing http://example.com instead of Google. If that works, the issue is site-specific—you'll need to tweak capabilities like User-Agent, enable cookies, or adjust JS settings to mimic a real browser more closely.
内容的提问来源于stack exchange,提问作者David C

