动态网页爬虫使用querySelectorAll返回不完整NodeList问题求助
问题原因
evaluate返回值存在序列化限制:Nightmare的evaluate方法运行在浏览器沙箱上下文,返回值必须是可序列化的JSON类型才能传递回Node.js环境。DOM节点、NodeList都属于浏览器原生对象,无法直接序列化,直接返回会丢失绝大部分属性,仅能保留少数可序列化的自定义属性。- 动态内容等待不足:目标站点的价格数据是前端异步请求后渲染的,
goto执行完成仅代表基础HTML加载完毕,此时5个span.amount节点还未完全渲染。
修复方案
你需要先等待目标节点全部渲染完成,再在evaluate内部提前提取需要的内容,组装成纯数组/普通对象后再返回,修改后代码如下:
const Nightmare = require('nightmare'); const nightmare = Nightmare({ show: true }); const url = 'https://mir4draco.com/price'; nightmare .goto(url) // 等待第5个span.amount节点渲染完成,确保所有目标节点已加载 .wait('span.amount:nth-of-type(5)') .evaluate(() => { // 手动提取节点内容组成纯数组,避免返回不可序列化的DOM对象 return Array.from(document.querySelectorAll('span.amount')).map(node => node.textContent.trim()); }) .end() .then(response => { console.log(response); // 输出包含5个价格文本的数组 }).catch(error => { console.error('Search failed:', error); });
如果需要获取节点的其他属性,可以在map遍历中手动提取为普通对象的字段即可:
return Array.from(document.querySelectorAll('span.amount')).map(node => ({ text: node.textContent.trim(), class: node.className, // 其他需要的属性都可在此处手动提取 }))
内容的提问来源于stack exchange,提问作者Matheus Ribeiro
相关产品推荐
相关产品推荐

