You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium和Python提取下载动态渲染图片?解决仅获首图等问题

解决动态渲染图片的提取与下载问题

Hey there, let's tackle your three main concerns one by one with practical, actionable solutions using Python and Selenium:

1. 为什么只能提取到第一个图片URL?

Chances are you're using driver.find_element() instead of driver.find_elements() to locate images. The former returns only the first matching element, while the latter gives you a list of all elements that match the selector.

Quick fix: Replace any instance of find_element with find_elements when targeting <img> tags. For example:

# 错误写法:只获取第一个图片元素
single_img = driver.find_element(By.TAG_NAME, 'img')
# 正确写法:获取页面所有图片元素
all_imgs = driver.find_elements(By.TAG_NAME, 'img')

Also, make sure the page has fully loaded all dynamic content before you scrape. Use explicit waits to ensure images are rendered properly:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 等待最多10秒,直到所有img元素可见
wait = WebDriverWait(driver, 10)
all_imgs = wait.until(EC.visibility_of_all_elements_located((By.TAG_NAME, 'img')))

2. 自动识别img标签,无需编写特定XPath

You don't need site-specific XPath at all! Just target the <img> tag directly using By.TAG_NAME, which works across all websites that use standard HTML image tags.

To cover cases where dynamic sites use lazy-loaded images (storing URLs in data-src instead of src), you can add a check for common alternative attributes:

def get_image_url(img_element):
    # 优先取懒加载属性,再回退到默认src
    url = img_element.get_attribute('data-src') or img_element.get_attribute('data-lazy-src') or img_element.get_attribute('src')
    return url

# 收集所有有效图片URL
image_urls = []
for img in all_imgs:
    url = get_image_url(img)
    if url:
        image_urls.append(url)

This approach automatically grabs all images without needing to tweak selectors for different sites.

3. 完整的Selenium + Python动态图片提取&下载流程

Here's an end-to-end script that handles dynamic rendering, waits for content, extracts images, and downloads them:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import requests
import os
from urllib.parse import urljoin

# 初始化浏览器驱动(确保你的ChromeDriver路径配置正确)
driver = webdriver.Chrome()
driver.get("YOUR_TARGET_URL")

# 等待页面加载完成,所有图片可见
wait = WebDriverWait(driver, 15)
all_imgs = wait.until(EC.visibility_of_all_elements_located((By.TAG_NAME, 'img')))

# 创建保存图片的文件夹
if not os.path.exists("downloaded_images"):
    os.makedirs("downloaded_images")

# 遍历图片,提取URL并下载
for idx, img in enumerate(all_imgs):
    img_url = img.get_attribute('data-src') or img.get_attribute('src')
    if not img_url:
        continue
    # 处理相对路径,转为绝对URL
    absolute_url = urljoin(driver.current_url, img_url)
    
    try:
        # 下载图片
        response = requests.get(absolute_url, stream=True)
        response.raise_for_status()
        
        # 保存图片
        with open(f"downloaded_images/image_{idx+1}.jpg", 'wb') as f:
            for chunk in response.iter_content(1024):
                f.write(chunk)
        print(f"下载成功:{absolute_url}")
    except Exception as e:
        print(f"下载失败 {absolute_url}:{str(e)}")

# 关闭浏览器
driver.quit()

Extra Tips for Tricky Cases:

  • Lazy-loaded images: If images load as you scroll, simulate scrolling to the bottom of the page to trigger all images to load:
    # 滚动到页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待新图片加载完成
    wait.until(EC.presence_of_all_elements_located((By.TAG_NAME, 'img')))
    
  • Images inside iframes: If images are nested in an iframe, switch to the iframe first before extracting:
    iframe = driver.find_element(By.TAG_NAME, 'iframe')
    driver.switch_to.frame(iframe)
    # 然后再执行图片提取逻辑
    
  • Anti-scraping measures: Add a user-agent, avoid rapid requests, or use a headless browser to mimic human behavior:
    options = webdriver.ChromeOptions()
    options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
    options.add_argument("--headless=new")  # 无头模式,不显示浏览器窗口
    driver = webdriver.Chrome(options=options)
    

内容的提问来源于stack exchange,提问作者venkat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:09:33