使用JSoup抓取网站图片遇128张限制,求突破方案
解决JSoup仅能抓取128张Feather图标的问题
问题根源很明确:目标网站采用动态加载机制,初始页面只渲染128张图标(对应data-page_size="128"),剩余图标会在用户滚动页面时通过AJAX请求加载。JSoup只能获取首次HTTP请求返回的静态HTML,无法捕获后续动态加载的内容。
下面是两种可行的解决方案:
方案1:直接调用网站的API接口(高效推荐)
网站在加载更多图标时会调用后端API接口,直接请求这些接口可以绕过页面渲染,快速获取所有图标数据:
- 打开浏览器开发者工具(F12),切换到「Network」标签,过滤「XHR」请求;
- 滚动页面加载更多图标,找到加载图标的API请求(通常是类似
/api/v4/icons的地址); - 分析请求参数,比如
family=feather(指定图标家族)、page(页码)、page_size(每页数量); - 直接通过代码循环请求API,解析返回的JSON数据提取图片URL。
示例代码(Java + JSoup + JSON解析):
import org.jsoup.Jsoup; import org.json.JSONArray; import org.json.JSONObject; import java.io.IOException; import java.io.InputStream; import java.net.URL; import java.nio.file.Files; import java.nio.file.Paths; public class FeatherIconDownloader { public static void main(String[] args) throws IOException { int page = 1; int total = 0; final String API_BASE = "https://www.iconfinder.com/api/v4/icons"; // 模拟浏览器UA,避免被拦截 final String USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"; while (true) { // 请求API,设置足够大的page_size减少请求次数 String jsonBody = Jsoup.connect(API_BASE) .userAgent(USER_AGENT) .data("family", "feather") .data("page", String.valueOf(page)) .data("page_size", "300") .ignoreContentType(true) // 允许解析JSON内容 .execute() .body(); JSONObject response = new JSONObject(jsonBody); JSONArray icons = response.getJSONArray("icons"); if (icons.length() == 0) break; // 无更多图标,终止循环 for (int i = 0; i < icons.length(); i++) { JSONObject icon = icons.getJSONObject(i); // 从JSON结构中提取PNG预览图URL(路径需根据实际返回调整) String pngUrl = icon.getJSONObject("raster_sizes") .getJSONArray("items").getJSONObject(0) .getJSONArray("formats").getJSONObject(0) .getString("preview_url"); // 下载图片到本地示例 String fileName = pngUrl.substring(pngUrl.lastIndexOf('/') + 1); try (InputStream in = new URL(pngUrl).openStream()) { Files.copy(in, Paths.get("./feather_icons/", fileName)); } total++; } page++; } System.out.println("完成,共下载 " + total + " 张图标"); } }
方案2:模拟浏览器滚动加载(更直观)
使用Selenium/Playwright等工具模拟浏览器行为,自动滚动页面触发动态加载,待所有图标加载完成后再抓取:
示例代码(Java + Selenium):
import org.openqa.selenium.By; import org.openqa.selenium.JavascriptExecutor; import org.openqa.selenium.WebDriver; import org.openqa.selenium.WebElement; import org.openqa.selenium.chrome.ChromeDriver; import java.util.List; import java.util.concurrent.TimeUnit; public class SeleniumIconScraper { public static void main(String[] args) { WebDriver driver = new ChromeDriver(); driver.get("https://www.iconfinder.com/search/icons?family=feather"); driver.manage().timeouts().implicitlyWait(10, TimeUnit.SECONDS); JavascriptExecutor js = (JavascriptExecutor) driver; int prevCount = 0; int currCount; // 循环滚动直到无新图标加载 do { prevCount = driver.findElements(By.cssSelector("div.continuous-pagination.icon-preview-inline img[src$=.png]")).size(); js.executeScript("window.scrollTo(0, document.body.scrollHeight);"); try { Thread.sleep(2000); } catch (InterruptedException e) { e.printStackTrace(); } currCount = driver.findElements(By.cssSelector("div.continuous-pagination.icon-preview-inline img[src$=.png]")).size(); } while (currCount > prevCount); // 提取所有图标URL并下载 List<WebElement> imgs = driver.findElements(By.cssSelector("div.continuous-pagination.icon-preview-inline img[src$=.png]")); System.out.println("共抓取到 " + imgs.size() + " 张图标"); for (WebElement img : imgs) { String src = img.getAttribute("src"); // 这里添加下载逻辑,参考方案1的下载代码 System.out.println(src); } driver.quit(); } }
注意事项
- 因为你有使用授权,仍需控制请求频率,避免给服务器造成过大压力;
- 若API请求被拦截,可添加请求头中的
Cookie(从浏览器请求中复制),或使用代理; - 定期检查API结构,网站可能会更新接口参数或返回格式。
内容的提问来源于stack exchange,提问作者olavSR
相关产品推荐
相关产品推荐

