You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium+Python爬取图片失败:无报错但文件夹为空

问题分析与修复方案

你的脚本存在几个关键问题,导致图片无法正常保存,逐一说明并修复:

1. 等待机制未实际生效

你创建了WebDriverWait对象但没有执行等待逻辑,页面还没加载完成就获取了page_source,此时eBay的商品图片可能还未渲染,BeautifulSoup自然抓不到目标元素。必须用WebDriverWait等待图片元素加载完成后再解析页面。

2. 图片保存路径错误

image.save()的第一个参数要求是完整文件路径(文件夹+文件名),你只传入了文件夹路径,直接导致保存操作失败。应该使用生成的唯一文件名拼接完整路径。

3. 元素去重逻辑错误

if name not in results中的name是<img>标签对象,而results列表存储的是图片URL字符串,两者永远不会匹配,会导致重复添加URL或逻辑失效。需要判断图片URL是否已在列表中。

4. 缺少请求头(反爬拦截)

直接用requests.get()请求图片,容易被eBay的反爬机制拦截,导致获取不到图片内容。需要添加模拟浏览器的请求头。

5. 函数代码缩进错误

gets_url函数内的代码没有缩进,这会触发语法错误(你说脚本无报错可能是粘贴格式问题),必须修正缩进保证函数逻辑正常执行。


修复后的完整代码

import hashlib
import io
import requests
from pathlib import Path
from PIL import Image
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver import ChromeOptions
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.wait import WebDriverWait

# 配置Chrome无头模式,添加UA模拟浏览器
options = ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
driver = webdriver.Chrome(options=options)

try:
    driver.get(
        "https://www.ebay.co.uk/b/bn_79257191?_trkparms=parentrq%3A17b9fd4e1910aaf4d205ce64ffff9878%7Cpageci%3Ae4c4c85d"
        "-5180-11ef-874e-06e2120eee8d%7Cc%3A1%7Ciid%3A2%7Cli%3A8874")
    
    # 等待商品图片容器加载完成,最多等待20秒
    wait = WebDriverWait(driver, 20)
    wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "s-item__image-wrapper")))
    
    # 获取渲染完成的页面源码
    content = driver.page_source
    soup = BeautifulSoup(content, "html.parser")
finally:
    # 确保浏览器进程被关闭,避免资源泄漏
    driver.quit()


def gets_url(classes, location, source):
    results = []
    for a in soup.findAll(attrs={"class": classes}):
        img_tag = a.find(location)
        if img_tag:
            img_url = img_tag.get(source)
            # 判断URL是否已存在,避免重复抓取
            if img_url not in results:
                results.append(img_url)
    return results


if __name__ == "__main__":
    # 创建图片保存目录,不存在则自动创建
    save_dir = Path("/home/kali/Documents/Images/")
    save_dir.mkdir(exist_ok=True)
    
    returned_results = gets_url("s-item__image-wrapper image-treatment", "img", "src")
    print(f"抓取到{len(returned_results)}张图片URL")
    
    # 请求图片的UA头,模拟浏览器
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    for idx, img_url in enumerate(returned_results):
        try:
            # 请求图片内容
            image_content = requests.get(img_url, headers=headers).content
            image_file = io.BytesIO(image_content)
            image = Image.open(image_file).convert("RGB")
            
            # 生成唯一文件名
            file_name = f"{hashlib.sha1(image_content).hexdigest()[:10]}.png"
            file_path = save_dir / file_name
            
            # 保存图片
            image.save(file_path, "PNG", quality=80)
            print(f"已保存图片: {file_path}")
        except Exception as e:
            print(f"保存图片失败 {img_url}: {str(e)}")

额外说明

  • 加入try...finally确保浏览器进程被正确关闭,避免资源泄漏。
  • 提前创建图片保存目录,避免因目录不存在导致保存失败。
  • 添加错误捕获,方便排查单张图片保存失败的原因。
  • 给Selenium和requests都添加了User-Agent,降低被反爬拦截的概率。

内容的提问来源于stack exchange,提问作者McBongPuff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 03:48:09