使用Selenium爬虫时遭遇Cloudflare 502 Bad Gateway错误求助
求助:Selenium爬取时频繁遭遇Cloudflare 502 Bad Gateway错误
我目前用Python的Selenium工具抓取某在线数据库,需要通过页面导航获取目标数据。但每次运行代码时,总会遇到Cloudflare的502 Bad Gateway错误,该错误有时会自行消失,但出现的时机似乎和循环执行的位置有关。恳请各位提供规避该错误的建议,以下是我与Chrome交互的代码片段:
# ! Final ! #### Define Driver & Starting URL #### # Location of chromedriver driver_path = "/Users/shrey/Desktop/Python Projects/Selenium/chromedriver" # Beginning url & initialize driver url = "https://tamu.libguides.com/az.php" driver = webdriver.Chrome() # Make driver wait for elements to load when find_element() is run for the rest of our code driver.implicitly_wait(10) # Launch driver driver.get(url) # Press "Ancestry Database" link driver.find_element(By.LINK_TEXT, "Ancestry Library").click() # Give time for user to login to database time.sleep(30) # Go to link where we can search from home = "https://www.ancestrylibrary.com/search/collections/1742/" driver.get(home) # Switch to first tab (Search tab we just opened) driver.switch_to.window(driver.window_handles[0]) #### Loop through each year present in the data #### for yr in range(1886, 1952): # Go to search home driver.get(home) # Find textbox & Input Year -------- year_input = driver.find_element(By.CSS_SELECTOR, "#sfs_SelfCivilYear") year_input.send_keys(str(yr)) # Press "search" button driver.find_element(By.CSS_SELECTOR, "#searchButton").click() # Determine number of times we need to loop -------- # Find text which includes total number of results (formatted as "Results 1–20 of 1,351") n_raw = driver.find_element(By.XPATH, '//*[@id="results-header"]/h3').text # Isolate the important number (1,351) n_num = (tot_results.split()[-1]) # pulls the last word from the string - our desired number # Remove comma and convert to number ("1,351" >>> 1351) n_total = int(re.sub(",", "", n_num)) # Determine number of loops we need to do to scrape all the data loop_count = math.floor(n_total/20) + 1 # Loop thru pages and collect links -------- # Init empty list links = [] # Loop n times (calc'd earlier) for i in range(loop_count): # If we are on our last iter, do the same but do not click "next page" button if i == range(loop_count)[-1]: # Find & Store all "View Result" links current_pg_links = driver.find_elements(By.CSS_SELECTOR, ".srchFoundDB a") # Loop through all links pulled & append for link in current_pg_links: # Get actual url from 'href' attribute url = link.get_attribute('href') # Append URL to final list links.append(url) else: # Find & Store all "View Result" links current_pg_links = driver.find_elements(By.CSS_SELECTOR, ".srchFoundDB a") for link in current_pg_links: # Get actual url from 'href' attribute url = link.get_attribute('href') # Append URL to final list links.append(url) # Press "next page" button driver.find_element(By.CSS_SELECTOR, "a.ancBtn.sml.green.icon.iconArrowRight").click()
规避502错误的实用建议
添加随机延迟,模拟真实用户节奏
循环执行速度过快会触发Cloudflare反爬,把固定的time.sleep()换成随机延迟,比如搜索、翻页后等待2-6秒:import random # 搜索后延迟 time.sleep(random.uniform(2, 6)) # 翻页后延迟 time.sleep(random.uniform(3, 7))替换隐式等待为显式等待
隐式等待全局生效,容易导致页面未加载完全就执行操作。改用WebDriverWait针对特定元素等待,确保元素就绪后再操作:from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 等待搜索按钮可点击 search_btn = WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.CSS_SELECTOR, "#searchButton")) ) search_btn.click()配置Chrome选项,隐藏自动化痕迹
默认ChromeDriver特征明显,添加用户代理、禁用自动化标识:from selenium.webdriver.chrome.options import Options chrome_options = Options() chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(options=chrome_options)添加错误重试机制
捕获502相关异常,重试当前操作,避免程序直接中断:from selenium.common.exceptions import WebDriverException max_retries = 3 for yr in range(1886, 1952): retry_count = 0 while retry_count < max_retries: try: driver.get(home) # 后续输入、搜索等操作 break except WebDriverException as e: if "502" in str(e): retry_count += 1 time.sleep(5) # 重试前等待5秒 else: raise减少重复页面请求
外层循环每次都调用driver.get(home),可以只在第一次进入或页面出错时重新加载,降低请求频率。
内容的提问来源于stack exchange,提问作者shrey_shankar
相关产品推荐
相关产品推荐

