Python如何处理Selenium爬虫访问空白页的异常以继续爬取后续URL?
修复方案
核心思路是补充超时限制、空白页检测逻辑,同时细化异常捕获规则,确保单URL异常不会中断整体爬取流程。
优化点说明
- 初始化driver时设置页面加载超时和隐式等待,避免页面无响应导致代码卡死
- 新增空白页校验:页面内容长度过短、无核心业务节点时直接判定为无效页跳过
- 用显式等待替代固定sleep,既提升爬取效率也减少元素未加载完成导致的误判
- 细化异常捕获类型,仅捕获爬虫相关的可跳过异常,避免掩盖严重代码错误
- 异常信息输出具体出错URL和原因,方便后续排查问题
修改后代码
from selenium import webdriver import time from bs4 import BeautifulSoup import pandas as pd from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, NoSuchElementException, WebDriverException # 初始化driver并配置全局超时 driver = webdriver.Chrome() driver.set_page_load_timeout(15) # 单页面最多加载15秒 driver.implicitly_wait(5) # 查找元素最多等待5秒 dataf=[] val=[] baseurl='https://careers.abbvie.com/' endurl='?lang=en-us&previousLocale=en-US' # 列表页采集逻辑加异常处理 for x in range(1,89): try: driver.get(f'https://careers.abbvie.com/abbvie/jobs?page={x}&categories=Administrative%20Services%7CBusiness%20Development%7CGeneral%20Management%7CHEOR%2FMarket%20Access%7CInformation%20Technology%7CMarketing%7CMedical%7CRegulatory%20Affairs%7CSales%7CSales%20Support') time.sleep(7) page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') eachRow = soup.find_all('p', class_='job-title') for link in eachRow: for links in link.find_all('a',href=True): val.append(baseurl+links['href']+endurl) print(f"第{x}页采集完成,当前累计URL数:{len(val)}") except Exception as e: print(f"第{x}页采集失败,跳过,错误:{str(e)}") continue # 详情页爬取逻辑优化 for b in val: try: driver.get(b) # 先判断是否空白页 page_source = driver.page_source if len(page_source.strip()) < 200: # 空白页内容长度通常极低,可根据实际情况调整阈值 print(f"检测到空白页,跳过URL:{b}") continue # 显式等待核心元素加载 wait = WebDriverWait(driver, 8) title = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="jibe-container"]/div[2]/div/div/h1'))).text location = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="header-locations"]/span'))).text categories = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="header-categories"]/span'))).text jobID = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="header-req_id"]/span'))).text job_dict = {"Title":title,"location":location,"categories":categories,"jobID":jobID,"URL":b} dataf.append(job_dict) print(f"爬取成功:{title}") except (TimeoutException, NoSuchElementException, WebDriverException) as e: print(f"爬取URL[{b}]失败,跳过,错误原因:{str(e)}") continue df=pd.DataFrame(dataf) df.to_csv('restasis.csv', encoding='utf-8-sig') # 加编码避免中文乱码 driver.quit()
补充说明
如果实际运行中发现空白页阈值不合适,可以调整len(page_source.strip()) < 200里的200数值,也可以额外增加核心节点判断,比如判断'jibe-container' not in page_source就判定为无效页。
内容的提问来源于stack exchange,提问作者Age
相关产品推荐
相关产品推荐

