You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用JSoup抓取网站图片遇128张限制,求突破方案

解决JSoup仅能抓取128张Feather图标的问题

问题根源很明确:目标网站采用动态加载机制,初始页面只渲染128张图标(对应data-page_size="128"),剩余图标会在用户滚动页面时通过AJAX请求加载。JSoup只能获取首次HTTP请求返回的静态HTML,无法捕获后续动态加载的内容。

下面是两种可行的解决方案:

方案1:直接调用网站的API接口(高效推荐)

网站在加载更多图标时会调用后端API接口,直接请求这些接口可以绕过页面渲染,快速获取所有图标数据:

  1. 打开浏览器开发者工具(F12),切换到「Network」标签,过滤「XHR」请求;
  2. 滚动页面加载更多图标,找到加载图标的API请求(通常是类似/api/v4/icons的地址);
  3. 分析请求参数,比如family=feather(指定图标家族)、page(页码)、page_size(每页数量);
  4. 直接通过代码循环请求API,解析返回的JSON数据提取图片URL。

示例代码(Java + JSoup + JSON解析):

import org.jsoup.Jsoup;
import org.json.JSONArray;
import org.json.JSONObject;

import java.io.IOException;
import java.io.InputStream;
import java.net.URL;
import java.nio.file.Files;
import java.nio.file.Paths;

public class FeatherIconDownloader {
    public static void main(String[] args) throws IOException {
        int page = 1;
        int total = 0;
        final String API_BASE = "https://www.iconfinder.com/api/v4/icons";
        // 模拟浏览器UA,避免被拦截
        final String USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36";

        while (true) {
            // 请求API,设置足够大的page_size减少请求次数
            String jsonBody = Jsoup.connect(API_BASE)
                    .userAgent(USER_AGENT)
                    .data("family", "feather")
                    .data("page", String.valueOf(page))
                    .data("page_size", "300")
                    .ignoreContentType(true) // 允许解析JSON内容
                    .execute()
                    .body();

            JSONObject response = new JSONObject(jsonBody);
            JSONArray icons = response.getJSONArray("icons");

            if (icons.length() == 0) break; // 无更多图标,终止循环

            for (int i = 0; i < icons.length(); i++) {
                JSONObject icon = icons.getJSONObject(i);
                // 从JSON结构中提取PNG预览图URL(路径需根据实际返回调整)
                String pngUrl = icon.getJSONObject("raster_sizes")
                        .getJSONArray("items").getJSONObject(0)
                        .getJSONArray("formats").getJSONObject(0)
                        .getString("preview_url");

                // 下载图片到本地示例
                String fileName = pngUrl.substring(pngUrl.lastIndexOf('/') + 1);
                try (InputStream in = new URL(pngUrl).openStream()) {
                    Files.copy(in, Paths.get("./feather_icons/", fileName));
                }
                total++;
            }
            page++;
        }
        System.out.println("完成,共下载 " + total + " 张图标");
    }
}

方案2:模拟浏览器滚动加载(更直观)

使用Selenium/Playwright等工具模拟浏览器行为,自动滚动页面触发动态加载,待所有图标加载完成后再抓取:

示例代码(Java + Selenium):

import org.openqa.selenium.By;
import org.openqa.selenium.JavascriptExecutor;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.chrome.ChromeDriver;

import java.util.List;
import java.util.concurrent.TimeUnit;

public class SeleniumIconScraper {
    public static void main(String[] args) {
        WebDriver driver = new ChromeDriver();
        driver.get("https://www.iconfinder.com/search/icons?family=feather");
        driver.manage().timeouts().implicitlyWait(10, TimeUnit.SECONDS);

        JavascriptExecutor js = (JavascriptExecutor) driver;
        int prevCount = 0;
        int currCount;

        // 循环滚动直到无新图标加载
        do {
            prevCount = driver.findElements(By.cssSelector("div.continuous-pagination.icon-preview-inline img[src$=.png]")).size();
            js.executeScript("window.scrollTo(0, document.body.scrollHeight);");
            try { Thread.sleep(2000); } catch (InterruptedException e) { e.printStackTrace(); }
            currCount = driver.findElements(By.cssSelector("div.continuous-pagination.icon-preview-inline img[src$=.png]")).size();
        } while (currCount > prevCount);

        // 提取所有图标URL并下载
        List<WebElement> imgs = driver.findElements(By.cssSelector("div.continuous-pagination.icon-preview-inline img[src$=.png]"));
        System.out.println("共抓取到 " + imgs.size() + " 张图标");
        for (WebElement img : imgs) {
            String src = img.getAttribute("src");
            // 这里添加下载逻辑,参考方案1的下载代码
            System.out.println(src);
        }

        driver.quit();
    }
}

注意事项

  • 因为你有使用授权,仍需控制请求频率,避免给服务器造成过大压力;
  • 若API请求被拦截,可添加请求头中的Cookie(从浏览器请求中复制),或使用代理;
  • 定期检查API结构,网站可能会更新接口参数或返回格式。

内容的提问来源于stack exchange,提问作者olavSR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 10:25:43