You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬虫时遭遇Cloudflare 502 Bad Gateway错误求助

求助:Selenium爬取时频繁遭遇Cloudflare 502 Bad Gateway错误

我目前用Python的Selenium工具抓取某在线数据库,需要通过页面导航获取目标数据。但每次运行代码时,总会遇到Cloudflare的502 Bad Gateway错误,该错误有时会自行消失,但出现的时机似乎和循环执行的位置有关。恳请各位提供规避该错误的建议,以下是我与Chrome交互的代码片段:

# ! Final !
#### Define Driver & Starting URL ####
# Location of chromedriver
driver_path = "/Users/shrey/Desktop/Python Projects/Selenium/chromedriver"

# Beginning url & initialize driver
url = "https://tamu.libguides.com/az.php"
driver = webdriver.Chrome()

# Make driver wait for elements to load when find_element() is run for the rest of our code
driver.implicitly_wait(10)

# Launch driver
driver.get(url)

# Press "Ancestry Database" link
driver.find_element(By.LINK_TEXT,
                    "Ancestry Library").click()

# Give time for user to login to database
time.sleep(30)

# Go to link where we can search from
home = "https://www.ancestrylibrary.com/search/collections/1742/"
driver.get(home)

# Switch to first tab (Search tab we just opened)
driver.switch_to.window(driver.window_handles[0])

#### Loop through each year present in the data ####
for yr in range(1886, 1952):
    # Go to search home
    driver.get(home)
    
    # Find textbox & Input Year --------
    year_input = driver.find_element(By.CSS_SELECTOR, "#sfs_SelfCivilYear")
    year_input.send_keys(str(yr))

    # Press "search" button
    driver.find_element(By.CSS_SELECTOR, "#searchButton").click()

    # Determine number of times we need to loop --------
    # Find text which includes total number of results (formatted as "Results 1–20 of 1,351")
    n_raw = driver.find_element(By.XPATH,
                                '//*[@id="results-header"]/h3').text

    # Isolate the important number (1,351)
    n_num = (tot_results.split()[-1]) # pulls the last word from the string - our desired number

    # Remove comma and convert to number ("1,351" >>> 1351)
    n_total = int(re.sub(",", "", n_num))

    # Determine number of loops we need to do to scrape all the data
    loop_count = math.floor(n_total/20) + 1

    # Loop thru pages and collect links --------
    # Init empty list
    links = []
    
    # Loop n times (calc'd earlier)
    for i in range(loop_count):
        
        # If we are on our last iter, do the same but do not click "next page" button
        if i == range(loop_count)[-1]: 
            # Find & Store all "View Result" links
            current_pg_links = driver.find_elements(By.CSS_SELECTOR, 
                                                    ".srchFoundDB a")

            # Loop through all links pulled & append
            for link in current_pg_links:
                # Get actual url from 'href' attribute
                url = link.get_attribute('href')

                # Append URL to final list
                links.append(url)

        else:
            # Find & Store all "View Result" links
            current_pg_links = driver.find_elements(By.CSS_SELECTOR, 
                                                    ".srchFoundDB a")

            for link in current_pg_links:
                # Get actual url from 'href' attribute
                url = link.get_attribute('href')

                # Append URL to final list
                links.append(url)

            # Press "next page" button
            driver.find_element(By.CSS_SELECTOR,
                                "a.ancBtn.sml.green.icon.iconArrowRight").click()

规避502错误的实用建议

  • 添加随机延迟,模拟真实用户节奏
    循环执行速度过快会触发Cloudflare反爬,把固定的time.sleep()换成随机延迟,比如搜索、翻页后等待2-6秒:

    import random
    # 搜索后延迟
    time.sleep(random.uniform(2, 6))
    # 翻页后延迟
    time.sleep(random.uniform(3, 7))
    
  • 替换隐式等待为显式等待
    隐式等待全局生效,容易导致页面未加载完全就执行操作。改用WebDriverWait针对特定元素等待,确保元素就绪后再操作:

    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    # 等待搜索按钮可点击
    search_btn = WebDriverWait(driver, 15).until(
        EC.element_to_be_clickable((By.CSS_SELECTOR, "#searchButton"))
    )
    search_btn.click()
    
  • 配置Chrome选项,隐藏自动化痕迹
    默认ChromeDriver特征明显,添加用户代理、禁用自动化标识:

    from selenium.webdriver.chrome.options import Options
    
    chrome_options = Options()
    chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
    chrome_options.add_argument("--disable-blink-features=AutomationControlled")
    chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
    chrome_options.add_experimental_option('useAutomationExtension', False)
    
    driver = webdriver.Chrome(options=chrome_options)
    
  • 添加错误重试机制
    捕获502相关异常,重试当前操作,避免程序直接中断:

    from selenium.common.exceptions import WebDriverException
    
    max_retries = 3
    for yr in range(1886, 1952):
        retry_count = 0
        while retry_count < max_retries:
            try:
                driver.get(home)
                # 后续输入、搜索等操作
                break
            except WebDriverException as e:
                if "502" in str(e):
                    retry_count += 1
                    time.sleep(5)  # 重试前等待5秒
                else:
                    raise
    
  • 减少重复页面请求
    外层循环每次都调用driver.get(home),可以只在第一次进入或页面出错时重新加载,降低请求频率。

内容的提问来源于stack exchange,提问作者shrey_shankar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 10:00:00