You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium批量爬取时随机出现页面无限加载无法访问问题

问题汇总

问题背景

  • 基于Python+Selenium编写简易爬虫,目标站点为食谱站点https://akispetretzikis.com/en,用于批量抓取食谱公开数据
  • 所有待爬取的食谱URL预先存储在Excel文件中,通过pandas读取后批量发起访问

异常现象

  • 爬取过程中会随机触发加载异常:计划爬取100条食谱时,for循环执行到随机索引位置就会卡住,页面进入无限加载状态无法正常打开
  • 故障触发点无固定规律:首次运行故障出现在索引i=21位置,将循环起始值设为20时故障出现在i=41位置,重新运行代码故障也可能出现在i=17位置,无固定复现路径

复现代码

def mainProgram(start):
    now = datetime.now()
    options = webdriver.ChromeOptions()
    options.add_argument("start-maximized")
    options.add_argument('--no-sandbox')
    options.add_argument('--disable-infobars')
    options.add_argument('--disable-dev-shm-usage')
    options.add_experimental_option('useAutomationExtension', False)
    options.add_argument('--disable-blink-features=AutomationControlled')                                                                         
    theDictionary = {"Link": [], "Name": [], "Time": [], "Difficulty": [],     
                     "Merides": [], "Ingredients": [],
                     "ThermidesPer100gr": [], "ThermidesAnaMerida": []}
    driver = webdriver.Chrome(executable_path=r'/usr/lib/chromium-browser/chromedriver', 
                              options=options)
    driver.set_window_size(1280, 960)                                                
    thePath = os.path.join(os.path.expanduser("~"), "Desktop", "ScrapeRecipes",   
                           "Cooking"+str(now.year)+".xlsx")
    thePathReadExcel = os.path.join(os.path.expanduser("~"), "Desktop", 
                                    "CookingUrls"+str(now.year)+".xlsx")
    UrlOfRecipes = readExcel(thePath=thePathReadExcel)


    try:
        Length = len(UrlOfRecipes)
        print(Length)
        Length = 100#e.g. 100 actual Length over 1k
        for i in range(start, Length, 1):
            driver.delete_all_cookies()
            driver.get(UrlOfRecipes["Link"][i])
            wait = WebDriverWait(driver, 20 + round(random.uniform(0, 4), 2))
            time.sleep(30 + round(random.uniform(0, 4), 2))  # mandatory sleep
            theDictionary["Link"].append(UrlOfRecipes["Link"][i])
            theDictionary = getDataFromRecipe(driver, theDictionary)
            time.sleep(20 + round(random.uniform(0, 4), 2))
            print(i)
    except Exception as e:
        print(e)
        writeOnExcel(theDict, thePath)
问题根因与修复方案

随机位置无限加载的问题不是网络波动导致,是默认Selenium配置缺陷+站点反爬拦截+代码逻辑bug三个因素共同导致的,和你写的长sleep逻辑没有直接关联。

  • 根因1:默认页面加载策略无兜底
    Selenium默认使用normal页面加载策略,必须等页面所有资源(含第三方广告、统计脚本、懒加载素材、外部字体)全部加载完成才会结束加载状态,只要任意一个第三方资源连接超时,就会一直卡在加载状态,这是无限加载的核心诱因。
    修复方式:
    1. 修改ChromeOptions配置,将页面加载策略改为eager,DOM树解析完成就停止等待,不需要等非核心资源加载:
      options.page_load_strategy = 'eager'
      
    2. 给driver设置硬超时阈值,单页加载超过阈值直接中断,避免进程卡死:
      driver.set_page_load_timeout(15)
      driver.set_script_timeout(10)
      
    3. 你代码中初始化了WebDriverWait但从未实际调用,等于没有显式等待逻辑,补全元素等待逻辑,等需要抓取的核心元素出现就执行数据提取,不要靠固定sleep等页面。
  • 根因2:单实例长期运行触发反爬+内存泄漏
    单个Chrome driver实例连续爬取几十个页面时,一方面浏览器内存占用会持续上涨触发内存泄漏,另一方面请求指纹、访问行为特征很容易被站点WAF识别,随机返回无限加载的拦截响应。
    修复方式:
    1. 不要全程只启动一次driver,每爬15-20个页面就关闭旧driver、重新启动新实例,彻底清空内存残留和指纹痕迹,不需要每次只删cookie。
    2. 把你现在单次30s+20s的超长固定sleep改成3-8秒的随机间隔,过长的固定停留反而会提升反爬识别概率。
    3. 页面加载完成后随机做1-2次滚动操作,模拟真人浏览行为,不要打开页面立刻提取数据。
  • 根因3:代码存在隐性bug
    异常捕获块中写入Excel调用的变量是theDict,但你实际存储数据的字典变量名为theDictionary,真触发异常时会因为变量不存在抛出二次错误,已爬取的数据无法正常落盘,需要把变量名统一。

内容的提问来源于stack exchange,提问作者gkasap

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 00:15:32