You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Puppeteer中使用XPath无法提取网页下一页链接该如何解决?

问题根因

你的代码存在两处核心错误,直接导致输出异常:

  • page.$x() 方法仅接收XPath表达式作为入参,不支持传入第二个回调函数直接读取DOM属性。该方法的返回值是匹配到的ElementHandle对象数组,你直接把这个对象和字符串拼接,就会输出JSHandle@node这类序列化后的异常内容。
  • page.waitForSelector(".company-listing") 前没有加await,会导致元素还没加载完成就执行后续XPath查询,有概率匹配不到目标元素。

修复后的代码

const puppeteer = require("puppeteer");
const base = "https://www.timesbusinessdirectory.com";
let url = "https://www.timesbusinessdirectory.com/company-listings";

(async () => {
    const browser = await puppeteer.launch({headless:false});
    const [page] = await browser.pages();
    await page.goto(url,{waitUntil: 'networkidle2'});
    // 加await确保元素加载完成
    await page.waitForSelector(".company-listing");
    // 先通过XPath拿到匹配的元素数组,取第一个匹配的下一页按钮
    const [nextBtnHandle] = await page.$x("//a[@aria-label='Next'][./span[@aria-hidden='true'][contains(.,'Next')]]");
    // 读取元素的href属性值
    const nextPageLink = await nextBtnHandle.evaluate(el => el.getAttribute('href'));
    url = base.concat(nextPageLink);
    console.log("========================>",url)
    await browser.close();
})();

运行上述代码即可得到预期输出:https://www.timesbusinessdirectory.com/company-listings?page=2

内容的提问来源于stack exchange,提问作者robots.txt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 07:27:04