如何用Selenium和Python提取下载动态渲染图片?解决仅获首图等问题
Hey there, let's tackle your three main concerns one by one with practical, actionable solutions using Python and Selenium:
1. 为什么只能提取到第一个图片URL?
Chances are you're using driver.find_element() instead of driver.find_elements() to locate images. The former returns only the first matching element, while the latter gives you a list of all elements that match the selector.
Quick fix: Replace any instance of find_element with find_elements when targeting <img> tags. For example:
# 错误写法:只获取第一个图片元素 single_img = driver.find_element(By.TAG_NAME, 'img') # 正确写法:获取页面所有图片元素 all_imgs = driver.find_elements(By.TAG_NAME, 'img')
Also, make sure the page has fully loaded all dynamic content before you scrape. Use explicit waits to ensure images are rendered properly:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 等待最多10秒,直到所有img元素可见 wait = WebDriverWait(driver, 10) all_imgs = wait.until(EC.visibility_of_all_elements_located((By.TAG_NAME, 'img')))
2. 自动识别img标签,无需编写特定XPath
You don't need site-specific XPath at all! Just target the <img> tag directly using By.TAG_NAME, which works across all websites that use standard HTML image tags.
To cover cases where dynamic sites use lazy-loaded images (storing URLs in data-src instead of src), you can add a check for common alternative attributes:
def get_image_url(img_element): # 优先取懒加载属性,再回退到默认src url = img_element.get_attribute('data-src') or img_element.get_attribute('data-lazy-src') or img_element.get_attribute('src') return url # 收集所有有效图片URL image_urls = [] for img in all_imgs: url = get_image_url(img) if url: image_urls.append(url)
This approach automatically grabs all images without needing to tweak selectors for different sites.
3. 完整的Selenium + Python动态图片提取&下载流程
Here's an end-to-end script that handles dynamic rendering, waits for content, extracts images, and downloads them:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import requests import os from urllib.parse import urljoin # 初始化浏览器驱动(确保你的ChromeDriver路径配置正确) driver = webdriver.Chrome() driver.get("YOUR_TARGET_URL") # 等待页面加载完成,所有图片可见 wait = WebDriverWait(driver, 15) all_imgs = wait.until(EC.visibility_of_all_elements_located((By.TAG_NAME, 'img'))) # 创建保存图片的文件夹 if not os.path.exists("downloaded_images"): os.makedirs("downloaded_images") # 遍历图片,提取URL并下载 for idx, img in enumerate(all_imgs): img_url = img.get_attribute('data-src') or img.get_attribute('src') if not img_url: continue # 处理相对路径,转为绝对URL absolute_url = urljoin(driver.current_url, img_url) try: # 下载图片 response = requests.get(absolute_url, stream=True) response.raise_for_status() # 保存图片 with open(f"downloaded_images/image_{idx+1}.jpg", 'wb') as f: for chunk in response.iter_content(1024): f.write(chunk) print(f"下载成功:{absolute_url}") except Exception as e: print(f"下载失败 {absolute_url}:{str(e)}") # 关闭浏览器 driver.quit()
Extra Tips for Tricky Cases:
- Lazy-loaded images: If images load as you scroll, simulate scrolling to the bottom of the page to trigger all images to load:
# 滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待新图片加载完成 wait.until(EC.presence_of_all_elements_located((By.TAG_NAME, 'img'))) - Images inside iframes: If images are nested in an iframe, switch to the iframe first before extracting:
iframe = driver.find_element(By.TAG_NAME, 'iframe') driver.switch_to.frame(iframe) # 然后再执行图片提取逻辑 - Anti-scraping measures: Add a user-agent, avoid rapid requests, or use a headless browser to mimic human behavior:
options = webdriver.ChromeOptions() options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") options.add_argument("--headless=new") # 无头模式,不显示浏览器窗口 driver = webdriver.Chrome(options=options)
内容的提问来源于stack exchange,提问作者venkat

