如何使用Puppeteer获取网页全部资源及CSS加载的图片?
嘿,这个问题问得相当精准——用Puppeteer抓取页面资源,确实得兼顾DOM里明面上的<img>标签,还有藏在CSS里的背景图这类“隐性”资源,我来一步步给你拆解清楚:
1. 先抓取DOM中的
<img>标签资源 这是最直接的部分,我们可以通过page.evaluate()在浏览器上下文里遍历所有<img>元素,提取它们的src属性。如果需要处理srcset(响应式图片)也能顺便扩展,但你这次聚焦图片,先搞定基础的:
const puppeteer = require('puppeteer'); (async () => { const browser = await puppeteer.launch(); const page = await browser.newPage(); // 等待网络空闲,确保页面资源加载完成 await page.goto('你的目标页面URL', { waitUntil: 'networkidle2' }); // 提取所有img标签的src(自动转为绝对URL) const domImageUrls = await page.evaluate(() => { return Array.from(document.querySelectorAll('img')).map(img => { // 把相对URL转成绝对URL,避免路径问题 return new URL(img.src, window.location.href).href; }); }); console.log('DOM中的图片资源:', domImageUrls); await browser.close(); })();
2. 抓取通过CSS规则加载的图片
这类图片包括background-image、list-style-image,甚至伪元素(:before/:after)里的背景图。我们需要遍历页面所有元素,获取它们的计算样式,从中提取图片URL:
// 接上面的代码,在page.goto之后添加: const cssImageUrls = await page.evaluate(() => { const uniqueUrls = new Set(); // 用Set自动去重 const allElements = document.querySelectorAll('*'); // 遍历页面所有元素 allElements.forEach(element => { const computedStyle = window.getComputedStyle(element); // 处理背景图 const bgImage = computedStyle.backgroundImage; if (bgImage && bgImage !== 'none') { // 匹配url()里的地址,兼容带引号/不带引号的情况 const urlMatches = bgImage.match(/url\(["']?(.*?)["']?\)/g); if (urlMatches) { urlMatches.forEach(match => { const rawUrl = match.replace(/url\(["']?(.*?)["']?\)/, '$1'); // 转成绝对URL const absoluteUrl = new URL(rawUrl, window.location.href).href; uniqueUrls.add(absoluteUrl); }); } } // 处理列表项图标 const listImage = computedStyle.listStyleImage; if (listImage && listImage !== 'none') { const urlMatch = listImage.match(/url\(["']?(.*?)["']?\)/); if (urlMatch) { const absoluteUrl = new URL(urlMatch[1], window.location.href).href; uniqueUrls.add(absoluteUrl); } } }); return Array.from(uniqueUrls); }); console.log('CSS加载的图片资源:', cssImageUrls);
3. 查看页面所有已加载的图片资源
如果想拿到页面实际加载过的所有图片(不管是通过DOM还是CSS,甚至是脚本动态加载的),可以监听Puppeteer的response事件,筛选出图片类型的响应:
const puppeteer = require('puppeteer'); (async () => { const browser = await puppeteer.launch(); const page = await browser.newPage(); const loadedImages = new Set(); // 监听所有网络响应,筛选图片资源 page.on('response', async (response) => { const url = response.url(); const contentType = response.headers()['content-type']; // 匹配所有图片MIME类型 if (contentType && contentType.startsWith('image/')) { loadedImages.add(url); } }); await page.goto('你的目标页面URL', { waitUntil: 'networkidle2' }); // 如果有懒加载图片,可以模拟滚动触发加载 // await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight)); // await page.waitForTimeout(1000); console.log('所有已加载的图片资源:', Array.from(loadedImages)); await browser.close(); })();
额外注意点
- 去重:上面的例子用
Set来避免重复资源,因为同一张图片可能同时通过DOM和CSS加载 - 动态资源:如果页面有懒加载(比如滚动才加载图片),可以模拟滚动或者等待特定元素加载完成
- 相对URL转绝对URL:用
new URL(rawUrl, window.location.href).href可以确保拿到的是完整可访问的地址
内容的提问来源于stack exchange,提问作者dsp_099
相关产品推荐
相关产品推荐

