You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Selenium的SKU图片自动搜索下载脚本开发求助

问题:Selenium定位图片XPath遭网站拦截,需完成SKU提取、图片下载流程

我需要编写自动化脚本实现以下功能:

  • 从Excel文件提取SKU码
  • 在指定网站搜索对应SKU
  • 获取商品图片URL并下载到本地,图片名称与SKU一致

目前已完成网站登录及SKU搜索输入操作,但定位图片XPath时遭遇网站拦截。现有Selenium代码如下:

from selenium import webdriver
from selenium.webdriver import Keys
from getpass import getpass
from selenium.webdriver.common.by import By

username = input("User: ")
password = getpass("Pass: ")

options = webdriver.ChromeOptions()
options.add_experimental_option("detach", True)
driver = webdriver.Chrome(options=options)

driver.get("http://portal.eaglebrandstrade.pt/#/")
driver.implicitly_wait(20)

username_textbox = driver.find_element(By.ID, "username")
username_textbox.send_keys(username)
password_textbox = driver.find_element(By.ID, "password")
password_textbox.send_keys(password)

button = driver.find_element(By.CLASS_NAME, "btn-primary")
button.click()

SKU = "lind-00024"
driver.get("http://portal.eaglebrandstrade.pt/#/skudetail")
driver.implicitly_wait(20)

mysearch = driver.find_element(By.XPATH, "//input[1]")
mysearch.send_keys(SKU)
mysearch.send_keys(Keys.ENTER)

driver.implicitly_wait(20)
img = driver.find_element(By.XPATH, "//img[1]")
img.get_attribute("srl")

解决思路与优化方案

1. 修复基础错误与优化元素定位

  • 属性拼写错误:代码中img.get_attribute("srl")应为img.get_attribute("src"),这是常见笔误。
  • 精准定位图片:避免使用//img[1]这种全局定位,结合页面结构用更具体的XPath,比如定位商品容器内的图片:
    # 示例:假设图片包裹在class为product-container的元素内
    img = driver.find_element(By.XPATH, "//div[contains(@class, 'product-container')]//img")
    img_url = img.get_attribute("src")
    

2. 绕过网站反爬拦截

网站可能通过检测自动化特征拦截请求,可添加以下配置模拟真实浏览器:

options = webdriver.ChromeOptions()
options.add_experimental_option("detach", True)
# 禁用自动化检测
options.add_argument("--disable-blink-features=AutomationControlled")
# 添加真实User-Agent
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 移除自动化标识
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)

driver = webdriver.Chrome(options=options)
# 执行JS进一步隐藏webdriver属性
driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")

3. 替换隐式等待为显式等待

隐式等待存在局限性,改用显式等待确保图片元素完全加载:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 等待图片元素出现,超时时间30秒
img = WebDriverWait(driver, 30).until(
    EC.presence_of_element_located((By.XPATH, "//div[contains(@class, 'product-image')]//img"))
)
img_url = img.get_attribute("src")

4. 补充Excel读取与图片下载功能

  • 读取Excel中的SKU:使用pandas快速提取SKU列表:
    import pandas as pd
    
    # 读取Excel,替换为你的文件路径和SKU列名
    df = pd.read_excel("sku_list.xlsx")
    sku_list = df["SKU"].tolist()
    
  • 下载图片到本地:使用requests库下载图片,命名为对应SKU:
    import requests
    
    def download_image(img_url, sku):
        try:
            # 添加请求头模拟浏览器访问
            headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"}
            response = requests.get(img_url, headers=headers, stream=True)
            response.raise_for_status()
            # 保存图片,格式可根据实际调整为png等
            with open(f"{sku}.jpg", "wb") as f:
                for chunk in response.iter_content(chunk_size=8192):
                    f.write(chunk)
            print(f"SKU {sku} 图片下载完成")
        except Exception as e:
            print(f"SKU {sku} 图片下载失败: {str(e)}")
    

内容的提问来源于stack exchange,提问作者Vicmol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 07:00:04