You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium图片爬取异常:脚本持续返回0图片链接求助

Google图片爬虫无法获取图片链接的排查与修复

问题描述

我正在开发一个小型项目,需要使用网络爬虫采集至多10000张图片。目前脚本运行无异常,但持续返回相同结果:Found: 0 image links, looking for more ...。想确认脚本是否存在被忽略的错误,或是有遗漏的配置?

原脚本

import os
import time
import requests
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

def fetch_image_urls(query, max_links_to_fetch, wd, sleep_between_interactions=1):
    def scroll_to_end(wd):
        wd.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(sleep_between_interactions)

    search_url = f"https://www.google.com/search?q={query}&tbm=isch"
    wd.get(search_url)

    image_urls = set()
    image_count = 0
    results_start = 0

    while image_count < max_links_to_fetch:
        scroll_to_end(wd)

        thumbnails = wd.find_elements(By.CSS_SELECTOR, "div[jsname='TVq1Bb'] img.rg_i")
        
        for img in thumbnails[results_start:]:
            try:
                img.click()
                time.sleep(sleep_between_interactions)
            except Exception as e:
                print(f"Error clicking image: {e}")
                continue

            large_image = wd.find_element(By.CSS_SELECTOR, "div[jsname='isl9D'] img.n3VNCb")
            image_url = large_image.get_attribute("src")

            if image_url and "http" in image_url:
                image_urls.add(image_url)

            image_count = len(image_urls)

            if len(image_urls) >= max_links_to_fetch:
                break
        else:
            print("Found:", len(image_urls), "image links, looking for more ...")
            time.sleep(30)
            continue
        
        results_start = len(thumbnails)

    return image_urls

def download_images(query, num_images, output_path):
    if not os.path.exists(output_path):
        os.makedirs(output_path)

    wd = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
    res = fetch_image_urls(query, num_images, wd=wd, sleep_between_interactions=1.5)

    for i, url in enumerate(res):
        try:
            image_content = requests.get(url).content
            with open(os.path.join(output_path, f'{query}_{i+1}.jpg'), 'wb') as f:
                f.write(image_content)
            print(f"Successfully downloaded {i+1}/{num_images} images")
        except Exception as e:
            print(f"Could not download {url} - {e}")

    wd.quit()

queries = ["medical imaging", "MRI scans", "CT scans", "X-ray images"]
for query in queries:
    download_images(query, num_images=500, output_path='medical_images')

运行结果

Found: 0 image links, looking for more ...
Found: 0 image links, looking for more ...
Found: 0 image links, looking for more ...
Found: 0 image links, looking for more ...
Found: 0 image links, looking for more ...
Found: 0 image links, looking for more ...
Found: 0 image links, looking for more ...
Found: 0 image links, looking for more ...

问题原因与修复方案

1. CSS选择器过时失效

Google图片页面的DOM结构会频繁更新,你使用的div[jsname='TVq1Bb'] img.rg_i(缩略图)和div[jsname='isl9D'] img.n3VNCb(大图)选择器已不再匹配当前页面元素。

修复:

  • 缩略图选择器替换为:img.Q4LuWd
  • 大图选择器替换为:img.sFlh5c.pT0Scc.iPVvYb

2. 未配置反爬策略,被Google识别为机器人

直接初始化的Selenium浏览器会带有自动化标识,容易被Google反爬机制拦截,导致页面无法正常加载图片或返回空结果。

修复:
初始化Chrome时添加反爬配置:

from selenium import webdriver

options = webdriver.ChromeOptions()
# 设置模拟浏览器的User-Agent
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 禁用自动化标识
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)

# 初始化浏览器时传入options
wd = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)

3. 等待逻辑不足,大图未加载完成就获取链接

单纯使用time.sleep无法保证大图完全加载,可能导致获取到空链接或错误元素。

修复:
使用Selenium的显式等待,等待大图元素加载完成后再获取链接:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException

# 替换原获取大图的代码
try:
    # 等待10秒,直到大图元素出现
    large_image = WebDriverWait(wd, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "img.sFlh5c.pT0Scc.iPVvYb"))
    )
    image_url = large_image.get_attribute("src")
except TimeoutException:
    print("大图加载超时,跳过当前图片")
    continue

4. 滚动逻辑未处理"加载更多"按钮

滚动到底部后,Google图片可能需要点击"加载更多"按钮才能获取更多结果,原脚本未处理此情况。

修复:
优化scroll_to_end函数,检查并点击加载更多按钮:

def scroll_to_end(wd):
    wd.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(sleep_between_interactions)
    # 检查是否存在加载更多按钮,存在则点击
    try:
        load_more_btn = wd.find_element(By.CSS_SELECTOR, "input.mye4qd")
        load_more_btn.click()
        time.sleep(sleep_between_interactions)
    except:
        # 没有加载更多按钮则跳过
        pass

内容的提问来源于stack exchange,提问作者Shaina Maisarah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 06:13:11