使用Puppeteer爬取玩家金币时出现TypeError错误求助
Puppeteer爬取plancke.io玩家BedWars金币报错及多用户查询实现
错误信息
C:\Users\Coach\Desktop\a\getuser.js:24 const txt = await el2.getProperty('textContent'); ^ TypeError: Cannot read properties of undefined (reading 'getProperty') at scrapeProduct (C:\Users\Coach\Desktop\a\getuser.js:24:27) at processTicksAndRejections (node:internal/process/task_queues:96:5) C:\Users\Coach\Desktop\a>PAUSE Press any key to continue . . .
原代码
console.clear(); const puppeteer = require('puppeteer'); console.log(` _______ _______ _______ ___ _______ __ __ _______ | || || _ || | | || |_| || | | || _ || |_| || | | ___|| || ___| | || | | || || | | |___ | || |___ | _|| |_| || || |___ ___ | ___| | | | ___| | |_ | || _ || || | | |___ | _ || |___ |_______||_______||__| |__||_______||___| |_______||__| |__||_______| `); async function scrapeProduct(url) { const browser = await puppeteer.launch(); const page = await browser.newPage(); await page.goto(url); const [el2] = await page.$x('//*[@id="collapse-1-2"]/div/ul/li[1]/text()'); const txt = await el2.getProperty('textContent'); const rawTxt = await txt.jsonValue(); console.log({rawTxt}) } scrapeProduct('https://plancke.io/hypixel/player/stats/coalwork#BedWars');
问题原因
- XPath文本节点处理错误:原代码XPath末尾加了
/text(),返回的是文本节点而非DOM元素,Puppeteer对文本节点的getProperty方法支持有限,导致el2为undefined。 - 缺少页面等待:plancke.io是动态渲染页面,直接访问后立即查询元素时,目标内容可能还未加载完成,导致定位失败。
- XPath路径适配问题:页面结构若有更新,原路径可能无法准确找到BedWars金币元素。
修复方案及多用户查询实现
单用户查询修复
修改代码逻辑,等待元素加载完成后再获取文本,避免文本节点的处理问题:
console.clear(); const puppeteer = require('puppeteer'); console.log(` _______ _______ _______ ___ _______ __ __ _______ | || || _ || | | || |_| || | | || _ || |_| || | | ___|| || ___| | || | | || || | | |___ | || |___ | _|| |_| || || |___ ___ | ___| | | | ___| | |_ | || _ || || | | |___ | _ || |___ |_______||_______||__| |__||_______||___| |_______||__| |__||_______| `); async function scrapeBedWarsCoins(username) { const browser = await puppeteer.launch({ headless: 'new' }); const page = await browser.newPage(); const url = `https://plancke.io/hypixel/player/stats/${username}#BedWars`; try { await page.goto(url, { waitUntil: 'networkidle2' }); // 等待目标元素加载完成 const [coinElement] = await page.waitForXPath('//*[@id="collapse-1-2"]/div/ul/li[1]', { timeout: 10000 }); if (!coinElement) { console.log(`用户${username}未找到BedWars数据`); await browser.close(); return; } // 通过evaluate获取元素文本内容 const coinText = await page.evaluate(el => el.textContent, coinElement); console.log(`用户${username}的BedWars金币: ${coinText.trim()}`); } catch (error) { console.log(`查询用户${username}时出错: ${error.message}`); } finally { await browser.close(); } } // 测试单用户查询 scrapeBedWarsCoins('coalwork');
多用户批量查询
扩展代码支持批量查询多个玩家,添加请求延迟避免反爬:
console.clear(); const puppeteer = require('puppeteer'); console.log(` _______ _______ _______ ___ _______ __ __ _______ | || || _ || | | || |_| || | | || _ || |_| || | | ___|| || ___| | || | | || || | | |___ | || |___ | _|| |_| || || |___ ___ | ___| | | | ___| | |_ | || _ || || | | |___ | _ || |___ |_______||_______||__| |__||_______||___| |_______||__| |__||_______| `); async function scrapeBedWarsCoins(username) { const browser = await puppeteer.launch({ headless: 'new' }); const page = await browser.newPage(); const url = `https://plancke.io/hypixel/player/stats/${username}#BedWars`; try { await page.goto(url, { waitUntil: 'networkidle2' }); const [coinElement] = await page.waitForXPath('//*[@id="collapse-1-2"]/div/ul/li[1]', { timeout: 10000 }); if (!coinElement) { console.log(`[${username}] 未找到BedWars金币数据`); return; } const coinText = await page.evaluate(el => el.textContent, coinElement); console.log(`[${username}] BedWars金币: ${coinText.trim()}`); } catch (error) { console.log(`[${username}] 查询失败: ${error.message}`); } finally { await browser.close(); } } // 批量处理用户列表 async function batchScrapeUsers(usernameList) { for (const username of usernameList) { await scrapeBedWarsCoins(username); // 添加1秒延迟,避免请求过于频繁触发反爬机制 await new Promise(resolve => setTimeout(resolve, 1000)); } } // 示例:传入多个玩家用户名查询 const targetUsers = ['coalwork', 'Player1', 'Player2']; batchScrapeUsers(targetUsers);
关键优化点
- 使用
page.waitForXPath替代直接page.$x,确保元素加载完成后再操作 - 移除XPath中的
/text(),直接定位DOM元素后通过evaluate获取文本 - 添加错误捕获和浏览器资源释放逻辑,避免内存泄漏
- 支持批量查询,添加请求延迟降低反爬风险
内容的提问来源于stack exchange,提问作者user20393634
相关产品推荐
相关产品推荐

