You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何处理Selenium爬虫访问空白页的异常以继续爬取后续URL?

修复方案

核心思路是补充超时限制、空白页检测逻辑,同时细化异常捕获规则,确保单URL异常不会中断整体爬取流程。

优化点说明

  • 初始化driver时设置页面加载超时和隐式等待,避免页面无响应导致代码卡死
  • 新增空白页校验:页面内容长度过短、无核心业务节点时直接判定为无效页跳过
  • 用显式等待替代固定sleep,既提升爬取效率也减少元素未加载完成导致的误判
  • 细化异常捕获类型,仅捕获爬虫相关的可跳过异常,避免掩盖严重代码错误
  • 异常信息输出具体出错URL和原因,方便后续排查问题

修改后代码

from selenium import webdriver
import time
from bs4 import BeautifulSoup
import pandas as pd
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, NoSuchElementException, WebDriverException

# 初始化driver并配置全局超时
driver = webdriver.Chrome()
driver.set_page_load_timeout(15)  # 单页面最多加载15秒
driver.implicitly_wait(5)  # 查找元素最多等待5秒

dataf=[]
val=[]
baseurl='https://careers.abbvie.com/'
endurl='?lang=en-us&previousLocale=en-US'

# 列表页采集逻辑加异常处理
for x in range(1,89):
    try:
        driver.get(f'https://careers.abbvie.com/abbvie/jobs?page={x}&categories=Administrative%20Services%7CBusiness%20Development%7CGeneral%20Management%7CHEOR%2FMarket%20Access%7CInformation%20Technology%7CMarketing%7CMedical%7CRegulatory%20Affairs%7CSales%7CSales%20Support')
        time.sleep(7)
        page_source = driver.page_source
        soup = BeautifulSoup(page_source, 'html.parser')
        eachRow = soup.find_all('p', class_='job-title')
        for link in eachRow:
            for links in link.find_all('a',href=True):
                val.append(baseurl+links['href']+endurl)
        print(f"第{x}页采集完成,当前累计URL数:{len(val)}")
    except Exception as e:
        print(f"第{x}页采集失败,跳过,错误:{str(e)}")
        continue

# 详情页爬取逻辑优化
for b in val:
    try:
        driver.get(b)
        # 先判断是否空白页
        page_source = driver.page_source
        if len(page_source.strip()) < 200: # 空白页内容长度通常极低,可根据实际情况调整阈值
            print(f"检测到空白页,跳过URL:{b}")
            continue
        # 显式等待核心元素加载
        wait = WebDriverWait(driver, 8)
        title = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="jibe-container"]/div[2]/div/div/h1'))).text
        location = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="header-locations"]/span'))).text
        categories = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="header-categories"]/span'))).text
        jobID = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="header-req_id"]/span'))).text
        job_dict = {"Title":title,"location":location,"categories":categories,"jobID":jobID,"URL":b}
        dataf.append(job_dict)
        print(f"爬取成功:{title}")
    except (TimeoutException, NoSuchElementException, WebDriverException) as e:
        print(f"爬取URL[{b}]失败,跳过,错误原因:{str(e)}")
        continue

df=pd.DataFrame(dataf)
df.to_csv('restasis.csv', encoding='utf-8-sig') # 加编码避免中文乱码
driver.quit()

补充说明

如果实际运行中发现空白页阈值不合适,可以调整len(page_source.strip()) < 200里的200数值,也可以额外增加核心节点判断,比如判断'jibe-container' not in page_source就判定为无效页。

内容的提问来源于stack exchange,提问作者Age

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 02:21:01