如何在Node.js+Puppeteer中提取li标签onclick里的完整URL?
从Puppeteer获取的onclick事件中提取完整跳转URL
方法1:直接提取onclick属性值(更高效)
没必要将onclick函数转为字符串,直接获取元素的onclick属性原始值,解析起来更简单。修改你的代码逻辑如下:
async function getUrls(){ const baseUrl = 'https://resale-intl.tickets.com/'; const getPartidos = await newPage.$$("li") for (const partido of getPartidos) { const onclickAttr = await newPage.evaluate((element) => { return element.getAttribute('onclick'); }, partido); // 用正则提取路径部分 const pathMatch = onclickAttr.match(/document.location.href = '([^']+)'/); if (pathMatch) { const relativePath = pathMatch[1].replace(/&/g, '&'); // 还原转义的&符号 const fullUrl = new URL(relativePath, baseUrl).href; console.log(fullUrl); } } } getUrls()
方法2:处理已有的onclick函数字符串
如果坚持要基于现有的函数toString结果处理,同样用正则匹配路径内容即可:
async function getUrls(){ const baseUrl = 'https://resale-intl.tickets.com/'; const getPartidos = await newPage.$$("li") for (const partido of getPartidos) { const urlPartidos = await newPage.evaluate((element) => { return element.onclick.toString() }, partido); // 匹配路径 const pathMatch = urlPartidos.match(/document.location.href = '([^']+)'/); if (pathMatch) { const relativePath = pathMatch[1].replace(/&/g, '&'); const fullUrl = new URL(relativePath, baseUrl).href; console.log(fullUrl); } } } getUrls()
关键细节说明
- 正则
/document.location.href = '([^']+)'/专门匹配单引号包裹的路径内容,[^']+确保只提取到单引号结束前的部分; - 使用
new URL(relativePath, baseUrl)拼接完整URL,比直接字符串拼接更可靠,能自动处理路径中的斜杠、相对路径等问题; - 替换
&为&是因为HTML中的转义规则,还原后才能保证URL参数的正确性。
内容的提问来源于stack exchange,提问作者Omar Giancarlo
相关产品推荐
相关产品推荐

