You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Puppeteer爬取PubChem页面href链接返回空值问题求助

Puppeteer爬取PubChem链接返回为空问题排查及解决方案

问题根因

  • 等待机制缺失:PubChem的搜索结果为客户端异步渲染,page.goto默认仅等待页面初始HTML加载完成,此时搜索请求尚未返回、目标元素未渲染到DOM中,直接查询自然返回空结果。
  • DOM查询结果处理错误:document.querySelectorAll返回的是NodeList类数组对象,原有代码直接将其整体塞入新数组[xxx],后续map遍历的是长度为1、仅包含NodeList对象的数组,而非NodeList内的a标签元素,无法正确读取href属性。
  • 选择器冗余脆弱:直接复制的浏览器全量选择器强依赖页面结构层级,前端只要调整元素嵌套关系就会失效。

修复后代码

const puppeteer = require('puppeteer')
puppeteer.launch({ headless: true }).then(async browser => {
    const page = await browser.newPage()
    await page.goto('https://pubchem.ncbi.nlm.nih.gov/#query=MES')
    // 等待目标元素渲染完成,最长超时设为10秒避免无限等待
    await page.waitForSelector('#featured-results .f-medium a', { timeout: 10000 })
    const links = await page.evaluate(() => {
        // 将NodeList转为真实数组后遍历取值
        return Array.from(document.querySelectorAll('#featured-results .f-medium a')).map(link => link.href)
    })
    links.forEach(link => console.log(link))
    
    await browser.close()
})

内容的提问来源于stack exchange,提问作者Walter Schrabmair

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 16:39:03