使用Python Selenium Chromedriver爬取Homes.com时频繁出现ERR_HTTP2_PROTOCOL_ERROR,但CURL请求可正常执行
Hey there, this is such a frustrating issue—especially when curl works but your Selenium setup doesn’t. The key clue here is that the request itself is valid, but the way Selenium-controlled Chrome is presenting itself or handling the connection is triggering this HTTP/2 error. Let’s walk through some actionable fixes you should try next:
1. Make your Selenium browser look less "automated"
Chrome sets subtle flags when it’s being controlled by Selenium, and anti-scraping systems pick up on these. Let’s mask those first:
from selenium import webdriver options = webdriver.ChromeOptions() # Disable the "automation controlled" flag that gives away Selenium control options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) # Match your regular Chrome's user agent (grab this from dev tools > console > navigator.userAgent) options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/134.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=options)
This tweak makes your Selenium instance way harder to distinguish from a manual browser session.
2. Inject the missing cache-control header you noticed
The successful curl request included cache-control: max-age=0, but your Selenium browser might not be sending this by default. Let’s add it using Chrome’s DevTools Protocol:
# Run this right after initializing the driver driver.execute_cdp_cmd('Network.setExtraHTTPHeaders', { "headers": { "cache-control": "max-age=0" } })
This ensures every request from your driver includes that header, matching the successful curl call you tested.
3. Reset your browser session periodically
Once you hit that first error, all subsequent requests fail—this means the server has flagged your current session. Try restarting the driver every few requests to get a fresh, unflagged session:
for idx, url in enumerate(urllist): # Restart driver every 5 requests to avoid persistent session flagging if idx % 5 == 0: if 'driver' in locals(): driver.quit() # Reinitialize driver with our custom options driver = webdriver.Chrome(options=options) driver.execute_cdp_cmd('Network.setExtraHTTPHeaders', {"headers": {"cache-control": "max-age=0"}}) delay = random.randint(2, 4) # Slightly longer, more natural delays than 1-3s time.sleep(delay) driver.get(url) # Your scraping logic here
You could also try incognito mode, but restarting the driver is more effective at fully resetting the server's view of you.
4. Force Chrome to use HTTP/1.1 instead of HTTP/2
Since the error is specific to HTTP/2, maybe the server is handling Selenium's HTTP/2 connections differently than curl's. Let’s disable HTTP/2 entirely:
options.add_argument("--disable-http2")
This makes your driver use HTTP/1.1, which might align better with how curl handles the request.
5. Add tiny, natural interactions to mimic real users
Even with delays, a browser that just loads pages without any interaction looks suspicious. Add small, random actions after each page load to blend in:
from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.common.by import By driver.get(url) # Randomly scroll a bit to mimic a user reading driver.execute_script(f"window.scrollTo(0, {random.randint(100, 600)});") # Move the mouse to a random element (like a div) to simulate natural movement elements = driver.find_elements(By.TAG_NAME, "div") if elements: ActionChains(driver).move_to_element(random.choice(elements)).perform() # Wait a tiny bit more before scraping to feel natural time.sleep(random.uniform(0.8, 1.5))
These small touches make your scraper look way more like a human browsing the site.
The core issue here is that Homes.com's anti-scraping measures are targeting automated browsers, but since your request works in curl, we just need to make your Selenium instance behave exactly like a real user. Start with masking the automation flags and adding the missing header—those two are the most likely fixes.
备注:内容来源于stack exchange,提问作者Ben

