You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Selenium Webdriver从动态网站下载图片

How to Scrape All Dynamically Loaded Images with Selenium in Java

Hey there! It sounds like you're dealing with lazy-loaded content—that's exactly why you're only grabbing ~40 image links instead of all of them. Most dynamic sites load content on-demand as the user scrolls down, so those extra image URLs don't exist in the page's DOM until their container enters the viewport. Let's walk through how to fix this step by step.

Step 1: Simulate Page Scrolling to Load All Content

The core fix here is to repeatedly scroll to the bottom of the page until no new content loads. We'll track the page height before and after each scroll to detect when we've hit the end of the content.

Here's a Java code snippet to implement this:

import org.openqa.selenium.JavascriptExecutor;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;

public class ScrollToLoadContent {
    public static void main(String[] args) {
        WebDriver driver = new ChromeDriver();
        driver.get("your-target-website-url");

        JavascriptExecutor js = (JavascriptExecutor) driver;
        long lastPageHeight = (long) js.executeScript("return document.body.scrollHeight");

        while (true) {
            // Scroll to the very bottom of the page
            js.executeScript("window.scrollTo(0, document.body.scrollHeight);");
            
            // Give the site time to load new content (adjust this based on site speed)
            try {
                Thread.sleep(2000);
            } catch (InterruptedException e) {
                e.printStackTrace();
            }

            long newPageHeight = (long) js.executeScript("return document.body.scrollHeight");
            if (newPageHeight == lastPageHeight) {
                // No new content loaded—we've reached the end
                break;
            }
            lastPageHeight = newPageHeight;
        }

        // Now all dynamically loaded content should be present in the DOM
    }
}

Step 2: Wait for All Image Elements to Fully Load

Even after scrolling, some images might still be in the process of loading. Use Selenium's explicit waits to ensure all <img> elements are present and their src attributes are populated:

import org.openqa.selenium.By;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;
import java.util.ArrayList;
import java.util.List;

// ... Add this inside your main method after scrolling ...
WebDriverWait wait = new WebDriverWait(driver, 10);
// Wait until all image elements are present on the page
wait.until(ExpectedConditions.presenceOfAllElementsLocatedBy(By.tagName("img")));

// Collect all valid image URLs
List<WebElement> imageElements = driver.findElements(By.tagName("img"));
List<String> allImageUrls = new ArrayList<>();
for (WebElement img : imageElements) {
    String imgSrc = img.getAttribute("src");
    // Skip empty or placeholder URLs if needed
    if (imgSrc != null && !imgSrc.isEmpty()) {
        allImageUrls.add(imgSrc);
    }
}

System.out.println("Total images found: " + allImageUrls.size());

Step 3: Download the Images

Once you have all the image URLs, you can use Java's built-in tools to download them. Here's a simple downloader method:

import java.io.FileOutputStream;
import java.io.InputStream;
import java.net.HttpURLConnection;
import java.net.URL;

public class ImageDownloader {
    public static void downloadImage(String imageUrl, String saveFilePath) {
        try {
            URL url = new URL(imageUrl);
            HttpURLConnection connection = (HttpURLConnection) url.openConnection();
            connection.setRequestMethod("GET");
            
            InputStream inputStream = connection.getInputStream();
            FileOutputStream outputStream = new FileOutputStream(saveFilePath);
            
            byte[] buffer = new byte[2048];
            int bytesRead;
            while ((bytesRead = inputStream.read(buffer)) != -1) {
                outputStream.write(buffer, 0, bytesRead);
            }
            
            outputStream.close();
            inputStream.close();
            System.out.println("Successfully downloaded: " + saveFilePath);
        } catch (Exception e) {
            System.err.println("Failed to download " + imageUrl + ": " + e.getMessage());
        }
    }
}

// Usage in your main method:
int imageCounter = 1;
for (String url : allImageUrls) {
    // Adjust the file extension based on the image type (e.g., .png, .webp)
    String savePath = "downloaded-images/image_" + imageCounter + ".jpg";
    ImageDownloader.downloadImage(url, savePath);
    imageCounter++;
}

Extra Tips for Reliability

  • Adjust Wait Times: If the site loads content slowly, increase the Thread.sleep duration or replace it with an explicit wait for a specific element that loads with new content (instead of fixed sleep).
  • Handle Edge Cases: Some sites load images only when elements are near the viewport (not just at the bottom). If scrolling to the bottom doesn't work, you can scroll to each unloaded image element individually.
  • Headless Mode: Run Chrome in headless mode to save resources if you don't need a visible browser:
    ChromeOptions options = new ChromeOptions();
    options.addArguments("--headless=new");
    WebDriver driver = new ChromeDriver(options);
    

内容的提问来源于stack exchange,提问作者Anurag Bhagsain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:40:56