Selenium图片爬取异常:脚本持续返回0图片链接求助
Google图片爬虫无法获取图片链接的排查与修复
问题描述
我正在开发一个小型项目,需要使用网络爬虫采集至多10000张图片。目前脚本运行无异常,但持续返回相同结果:Found: 0 image links, looking for more ...。想确认脚本是否存在被忽略的错误,或是有遗漏的配置?
原脚本
import os import time import requests from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager def fetch_image_urls(query, max_links_to_fetch, wd, sleep_between_interactions=1): def scroll_to_end(wd): wd.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(sleep_between_interactions) search_url = f"https://www.google.com/search?q={query}&tbm=isch" wd.get(search_url) image_urls = set() image_count = 0 results_start = 0 while image_count < max_links_to_fetch: scroll_to_end(wd) thumbnails = wd.find_elements(By.CSS_SELECTOR, "div[jsname='TVq1Bb'] img.rg_i") for img in thumbnails[results_start:]: try: img.click() time.sleep(sleep_between_interactions) except Exception as e: print(f"Error clicking image: {e}") continue large_image = wd.find_element(By.CSS_SELECTOR, "div[jsname='isl9D'] img.n3VNCb") image_url = large_image.get_attribute("src") if image_url and "http" in image_url: image_urls.add(image_url) image_count = len(image_urls) if len(image_urls) >= max_links_to_fetch: break else: print("Found:", len(image_urls), "image links, looking for more ...") time.sleep(30) continue results_start = len(thumbnails) return image_urls def download_images(query, num_images, output_path): if not os.path.exists(output_path): os.makedirs(output_path) wd = webdriver.Chrome(service=Service(ChromeDriverManager().install())) res = fetch_image_urls(query, num_images, wd=wd, sleep_between_interactions=1.5) for i, url in enumerate(res): try: image_content = requests.get(url).content with open(os.path.join(output_path, f'{query}_{i+1}.jpg'), 'wb') as f: f.write(image_content) print(f"Successfully downloaded {i+1}/{num_images} images") except Exception as e: print(f"Could not download {url} - {e}") wd.quit() queries = ["medical imaging", "MRI scans", "CT scans", "X-ray images"] for query in queries: download_images(query, num_images=500, output_path='medical_images')
运行结果
Found: 0 image links, looking for more ... Found: 0 image links, looking for more ... Found: 0 image links, looking for more ... Found: 0 image links, looking for more ... Found: 0 image links, looking for more ... Found: 0 image links, looking for more ... Found: 0 image links, looking for more ... Found: 0 image links, looking for more ...
问题原因与修复方案
1. CSS选择器过时失效
Google图片页面的DOM结构会频繁更新,你使用的div[jsname='TVq1Bb'] img.rg_i(缩略图)和div[jsname='isl9D'] img.n3VNCb(大图)选择器已不再匹配当前页面元素。
修复:
- 缩略图选择器替换为:
img.Q4LuWd - 大图选择器替换为:
img.sFlh5c.pT0Scc.iPVvYb
2. 未配置反爬策略,被Google识别为机器人
直接初始化的Selenium浏览器会带有自动化标识,容易被Google反爬机制拦截,导致页面无法正常加载图片或返回空结果。
修复:
初始化Chrome时添加反爬配置:
from selenium import webdriver options = webdriver.ChromeOptions() # 设置模拟浏览器的User-Agent options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 禁用自动化标识 options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) # 初始化浏览器时传入options wd = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
3. 等待逻辑不足,大图未加载完成就获取链接
单纯使用time.sleep无法保证大图完全加载,可能导致获取到空链接或错误元素。
修复:
使用Selenium的显式等待,等待大图元素加载完成后再获取链接:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException # 替换原获取大图的代码 try: # 等待10秒,直到大图元素出现 large_image = WebDriverWait(wd, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "img.sFlh5c.pT0Scc.iPVvYb")) ) image_url = large_image.get_attribute("src") except TimeoutException: print("大图加载超时,跳过当前图片") continue
4. 滚动逻辑未处理"加载更多"按钮
滚动到底部后,Google图片可能需要点击"加载更多"按钮才能获取更多结果,原脚本未处理此情况。
修复:
优化scroll_to_end函数,检查并点击加载更多按钮:
def scroll_to_end(wd): wd.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(sleep_between_interactions) # 检查是否存在加载更多按钮,存在则点击 try: load_more_btn = wd.find_element(By.CSS_SELECTOR, "input.mye4qd") load_more_btn.click() time.sleep(sleep_between_interactions) except: # 没有加载更多按钮则跳过 pass
内容的提问来源于stack exchange,提问作者Shaina Maisarah
相关产品推荐
相关产品推荐

