You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中WebDriver打开链接异常:跳转下一个链接无响应问题求助

WebDriver调用driver.get()仅填入URL不跳转,程序冻结的解决办法

问题重现

我写了一段Python脚本,用来批量抓取网页图片:读取本地的链接列表,逐个打开链接、提取图片并下载到桌面,然后自动跳转到下一个链接。但运行时遇到了奇怪的问题:当尝试跳转到下一个链接时,浏览器地址栏确实填入了新的URL,但就是不执行跳转动作,程序也没有报错,直接卡住不动了。

原代码如下:

import time
from selenium.common.exceptions import NoSuchElementException
import requests

# Preparar variáveis
time.sleep(1)
lines = [line.rstrip('\n') for line in open("C:/Users/Luís/Desktop/PEBTXT/linkdasimagens.txt")]
x = 0

# Loop para abrir imagens
while True:
    try:
        time.sleep(3)
        y = lines[x]
        driver.get(y)
        # Pegar imagem
        img = driver.find_element_by_xpath("//img[@class='BRnoselect']")
        src = img.get_attribute('src')
        with open(("C:/Users/Luís/Desktop/PEBTXT/file{}.jpg".format(x)), "wb") as f:
            f.write(requests.get(src).content)
        print(x)
        x += 1
        print(x)
    except NoSuchElementException:
        pass

解决方法

经过调试,我发现通过重复调用driver.get()方法,并在两次调用之间添加短暂等待,就能解决这个跳转失效的问题。修改后的关键代码如下:

time.sleep(2)
y = lines[x]
driver.get(y)
time.sleep(1)
driver.get(y)

原因分析

这种现象一般和WebDriver与浏览器的异步交互有关:

  • 第一次调用driver.get()时,可能因为浏览器进程未完全响应、前一个页面的资源未释放,或者目标网站的加载机制导致URL仅被填入地址栏,却没有触发导航动作
  • 添加短暂等待后再次调用driver.get(),相当于重新触发导航指令,确保浏览器执行跳转
  • 也有可能是目标网站存在某种前端拦截逻辑,单次请求未被正确处理,重复请求可以绕过这种限制

优化建议

为了让脚本更健壮,这里有几个改进点:

  • 替换固定等待为显式等待:time.sleep()是固定时长等待,不够灵活。推荐使用Selenium的WebDriverWait,根据页面元素的加载状态来动态等待,比如等待目标图片元素出现后再进行抓取:
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 等待图片元素加载完成,最多等待10秒
img = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, "//img[@class='BRnoselect']"))
)
  • 添加循环边界判断:原代码的while True会无限循环,当x超过链接列表长度时会抛出IndexError。建议改为:
while x < len(lines):
    # 原循环逻辑
  • 捕获更多异常:除了NoSuchElementException,还应该捕获requests的网络异常、文件IO异常等,避免程序意外崩溃:
import requests.exceptions

try:
    # 原代码逻辑
    f.write(requests.get(src).content)
except requests.exceptions.RequestException as e:
    print(f"下载图片失败: {e}")
except IOError as e:
    print(f"保存文件失败: {e}")

内容的提问来源于stack exchange,提问作者Luís Henrique Martins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:54:31