You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Puppeteer中使用XPath抓取下一页链接失败该如何解决

解决Puppeteer中XPath无法提取下一页链接的问题

错误原因

page.$x()的调用规则和CSS选择器对应的page.$eval()不同:page.$x()仅接收XPath表达式作为入参,返回值是匹配到的元素句柄数组,不支持直接传入回调函数提取属性,你原来的写法传了第二个回调参数属于无效用法。

修正后的实现代码

const puppeteer = require("puppeteer");
let url = "https://stackoverflow.com/questions/tagged/web-scraping";
 
(async () => {
    const browser = await puppeteer.launch({headless:false});
    const [page] = await browser.pages();
    
    await page.goto(url,{waitUntil: 'networkidle2'});
    // 用$x取匹配rel=next的a标签,解构取第一个匹配元素的句柄
    const [nextPageBtn] = await page.$x("//a[@rel='next']");
    // 调用元素句柄的evaluate方法提取href属性
    const nextPageLink = await nextPageBtn.evaluate(elem => elem.href);
    console.log("next page:",nextPageLink);
    await browser.close();
})();

补充注意事项

  • 实际生产环境建议增加非空判断,避免页面没有下一页按钮时报错:
    if (nextPageBtn) {
      const nextPageLink = await nextPageBtn.evaluate(elem => elem.href);
      console.log("next page:",nextPageLink);
    }
    
  • 如果你需要获取标签上的相对路径href,而非拼接后的完整链接,可以把evaluate内的返回值改为elem.getAttribute('href')即可。

内容的提问来源于stack exchange,提问作者MITHU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 23:57:04