使用Puppeteer与Cheerio抓取卡片列表时数据缺失的解决方法
Wallapop商品卡片抓取数据缺失问题解决
我尝试抓取Wallapop平台的商品卡片页面,需提取卡片的标题、价格、图片源及其他属性,但使用Puppeteer和Cheerio抓取时部分数据缺失(见截图)。请问如何确保所有数据正常抓取?
我的代码如下:
(async () => { try { const StealthPlugin = require("puppeteer-extra-plugin-stealth"); puppeteer2.use(StealthPlugin()); const browser = await puppeteer2.launch({ executablePath: "/usr/bin/chromium-browser", headless: true, args: [ "--no-sandbox", "--disable-setuid-sandbox", "--user-agent=" + USER_AGENT + "", ], }); const page = await browser.newPage({ignoreHTTPSErrors: true}); await page.setDefaultNavigationTimeout(0); await page.goto("https://es.wallapop.com/search?keywords=", { waitUntil: "networkidle0", }); await page.waitForTimeout(30000); const body = await page.evaluate(() => { return document.querySelector("body").innerHTML; }); var $ = cheerio.load(body); const pageItems = $(".ItemCardList__item .ng-star-inserted") .toArray() .map((item) => { const $item = $(item); return { // id: $item.attr('data-adid'), c10420p([^i]*)/ id: uuid.v4(), title: $item.find(".ItemCard__info").text(), link: "https://es.wallapop.com/item/", image: $item.find(".w-100").attr("src"), price: $item .find(".ItemCard__price") .text() .replace(/[_\W]+/g, ""), empresa: "wallapop", }; }); const allItems = items.concat(pageItems); console.log( pageItems.length, "items retrieved", allItems.length, "acumulat ed", ); // ... })
解决方案
1. 优化页面等待逻辑
- 替换固定时长等待,改为等待目标商品卡片元素加载完成,确保页面渲染完毕:
// 等待商品卡片容器出现,超时时间设为30秒 await page.waitForSelector(".ItemCardList__item .ng-star-inserted", { timeout: 30000 }); - 可以调整
goto的waitUntil参数为networkidle2,平衡等待效率和渲染完整性:await page.goto("https://es.wallapop.com/search?keywords=", { waitUntil: "networkidle2" });
2. 修正元素选择器与属性提取
- 标题选择器可能不准确,实际标题通常在
.ItemCard__title子元素中,需精准定位并去除空白字符:title: $item.find(".ItemCard__title").text().trim(), - 图片可能采用懒加载,实际地址存储在
data-src属性中,需优先读取该属性:image: $item.find(".w-100").attr("data-src") || $item.find(".w-100").attr("src"), - 商品链接需要拼接真实ID,可从元素的
data-adid属性获取:const adId = $item.attr("data-adid"); link: adId ? `https://es.wallapop.com/item/${adId}` : "",
3. 直接在Puppeteer中提取数据
避免将整个HTML传递给Cheerio,直接在page.evaluate内提取数据,确保获取JS渲染后的完整内容:
const pageItems = await page.evaluate(() => { return Array.from(document.querySelectorAll(".ItemCardList__item .ng-star-inserted")).map(item => { const titleEl = item.querySelector(".ItemCard__title"); const priceEl = item.querySelector(".ItemCard__price"); const imageEl = item.querySelector(".w-100"); const adId = item.getAttribute("data-adid"); return { id: adId || uuid.v4(), title: titleEl ? titleEl.textContent.trim() : "", link: adId ? `https://es.wallapop.com/item/${adId}` : "", image: imageEl ? (imageEl.getAttribute("data-src") || imageEl.getAttribute("src")) : "", price: priceEl ? priceEl.textContent.replace(/[_\W]+/g, "") : "", empresa: "wallapop" }; }); });
4. 增强反爬规避
- 添加视口设置,模拟真实浏览器窗口:
await page.setViewport({ width: 1280, height: 720 }); - 随机化等待时间,避免固定延迟被识别:
await page.waitForTimeout(Math.floor(Math.random() * 5000) + 2000); - 若存在滚动加载更多商品,模拟页面滚动触发加载:
await page.evaluate(async () => { await new Promise((resolve) => { let totalHeight = 0; const distance = 100; const timer = setInterval(() => { const scrollHeight = document.body.scrollHeight; window.scrollBy(0, distance); totalHeight += distance; if (totalHeight >= scrollHeight) { clearInterval(timer); resolve(); } }, 100); }); });
内容的提问来源于stack exchange,提问作者josiland
相关产品推荐
相关产品推荐

