使用Selenium爬取谷歌图片遇多问题:AttributeError及图片下载求助
谷歌图片爬取问题修复方案
问题汇总
- 报错
AttributeError: 'list' object has no attribute 'timeout' - 无法处理并下载base64格式的图片
download_image函数存在功能缺陷- 不清楚如何下载普通URL格式的图片资源
错误原因分析与修复
1. AttributeError 错误原因
代码中调用 download_image(img_link, file_name) 时,传入的 img_link 是通过 EC.presence_of_all_elements_located 获取的图片元素列表,而不是单个图片的URL。urllib.request.urlopen 接收的参数应该是字符串格式的URL,传入列表就会触发该错误。
2. Base64图片处理缺失
谷歌图片搜索结果中,部分图片的src是base64编码的格式(以data:image/开头),直接用urlopen无法处理,需要解码后再保存。
3. download_image 函数优化
原函数仅支持普通URL下载,缺少异常处理和Base64格式支持,需要重构以兼容两种场景。
4. 普通图片URL下载
确保传入正确的URL字符串,同时添加异常捕获,避免因网络问题或无效URL导致程序崩溃。
修复后的完整代码
from urllib.parse import urlparse from selenium import webdriver import time as t from selenium.webdriver.common.keys import Keys from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import urllib import base64 import os from urllib.error import URLError, HTTPError # 创建根目录 try: root_dir = "G:/Smokking_Project" os.makedirs(root_dir, exist_ok=True) except Exception as e: print(f"创建根目录失败: {e}") name = "smoked" # 创建分类子目录 sub_dir = os.path.join(root_dir, name) try: os.makedirs(sub_dir, exist_ok=True) except Exception as e: print(f"创建子目录失败: {e}") chrome_options = webdriver.ChromeOptions() chrome_options.add_experimental_option("excludeSwitches", ['enable-automation']) driver = webdriver.Chrome(options=chrome_options) wait = WebDriverWait(driver, 5) search_url = "https://www.google.com/search?q=smokinng&tbm=isch&ved=2ahUKEwi8k9zn9eOBAxVtlycCHTa_DnUQ2-cCegQIABAA&oq=smokinng&gs_lcp=CgNpbWcQAzIJCAAQGBCABBAKMgkIABAYEIAEEAoyCQgAEBgQgAQQCjoECCMQJzoFCAAQgAQ6BggAEAUQHjoECAAQHjoICAAQgAQQsQM6BAgAEAM6BwgAEBgQgARQjwdY8xJg-RloAHAAeACAAb0BiAHsCZIBAzAuOZgBAKABAaoBC2d3cy13aXotaW1nwAEB&sclient=img&ei=uUwhZfzSFO2unsEPtv66qAc&bih=723&biw=1517&hl=en" driver.get(search_url) t.sleep(3) links = [] x = 1 last_height = 0 def download_image(url, filename): try: # 处理Base64格式图片 if url.startswith("data:image/"): # 分割base64数据部分 base64_data = url.split(",")[1] img_bytes = base64.b64decode(base64_data) with open(filename, "wb") as output: output.write(img_bytes) print(f"Base64图片保存成功: {filename}") # 处理普通URL图片 else: resource = urllib.request.urlopen(url, timeout=10) with open(filename, "wb") as output: output.write(resource.read()) print(f"URL图片保存成功: {filename}") except (URLError, HTTPError, ValueError) as e: print(f"下载失败 {url}: {e}") except Exception as e: print(f"未知错误 {url}: {e}") while True: # 滚动到底部加载更多图片 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") t.sleep(4) # 获取所有图片元素 img_elements = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//a[1]/div[1]/img'))) t.sleep(1) for img in img_elements: url = img.get_attribute('src') if url and url not in links: links.append(url) print(f"发现图片链接: {url}") # 生成文件名 file_name = os.path.join(sub_dir, f"{x}.jpg") # 调用下载函数,传入正确的URL download_image(url, file_name) x += 1 # 检查是否已加载完所有内容 new_height = driver.execute_script("return document.body.scrollHeight") print(f"当前页面高度: {new_height}") if new_height == last_height: break last_height = new_height driver.close()
关键修复点说明
- 修复传参错误:将
download_image(img_link, file_name)改为download_image(url, file_name),传入单个图片的URL字符串。 - 添加Base64处理:通过判断URL前缀识别Base64格式,解码后保存为图片文件。
- 优化目录创建:使用
os.makedirs(exist_ok=True)避免重复创建目录的冗余代码,用os.path.join处理路径适配不同系统。 - 增加异常处理:在下载函数中捕获网络错误、解码错误等,避免程序中途崩溃。
- 导入缺失模块:添加
os和urllib.error相关模块的导入,解决原代码中未导入os导致的报错。
内容的提问来源于stack exchange,提问作者Ahmed Gaber
相关产品推荐
相关产品推荐

