使用Puppeteer抓取Looker Studio动态报表数据遇空内容问题求助
解决Puppeteer抓取Looker Studio动态数据为空的问题
问题分析
你遇到的空内容问题,核心原因是Looker Studio的报表内容嵌套在iframe中,直接获取document.body.innerText只能拿到外层页面内容,无法访问iframe内部的动态数据;其次旧版无头模式可能触发网站反爬检测,导致内容无法正常渲染。
修改方案
以下是针对性优化后的脚本:
import puppeteer from 'puppeteer'; async function fetchData() { try { const url = 'https://lookerstudio.google.com/u/0/reporting/e36054dd-ffc0-4ef4-b8ab-4d10f7ab4cda/page/wmP0D'; const options = { args: [ '--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage', '--disable-accelerated-2d-canvas', '--no-first-run', '--disable-gpu' ], // 使用新版无头模式,更贴近真实浏览器行为,规避反爬 headless: "new" }; const browser = await puppeteer.launch(options); const page = await browser.newPage(); await page.setUserAgent('Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'); await page.setViewport({width: 1920, height: 1080}); // 移除请求拦截,避免阻止关键JS资源加载(若需优化性能,可仅拦截非必要的大体积资源) // await page.setRequestInterception(true); // page.on('request', (req) => { // if (req.resourceType() === 'stylesheet' || req.resourceType() === 'font' || req.resourceType() === 'image') { // req.abort(); // } else { // req.continue(); // } // }); await page.goto(url, {waitUntil: 'networkidle2', timeout: 60000}); // 定位报表所在iframe并切换到其上下文 const iframeElement = await page.waitForSelector('iframe[src*="reporting"]', {timeout: 60000}); const frame = await iframeElement.contentFrame(); // 等待表格元素加载(需根据实际页面HTML结构调整选择器) await frame.waitForSelector('table', {timeout: 60000}); // 提取表格前3行数据(可按需调整行数) const tableData = await frame.evaluate(() => { const rows = document.querySelectorAll('tr'); const data = []; for (let i = 0; i < Math.min(3, rows.length); i++) { const cells = rows[i].querySelectorAll('td, th'); data.push(Array.from(cells).map(cell => cell.innerText.trim())); } return data; }); console.log('表格前几行数据:', tableData); await page.close(); await browser.close(); } catch (error) { console.error('Error fetching data:', error); } } fetchData();
关键调整说明
- 切换新版无头模式:将
headless: true改为headless: "new",新版模式更接近真实Chrome环境,降低被反爬识别的概率。 - 处理iframe嵌套:Looker Studio报表内容默认放在iframe中,必须先定位iframe并切换到其上下文,才能访问内部动态数据。
- 移除不必要的请求拦截:原拦截规则可能影响关键JS资源加载,导致动态内容无法渲染,暂时注释以确保页面完整加载。
- 精准等待目标元素:直接等待表格元素而非外层容器,确保数据已完全渲染完成。
- 定向提取数据:不再获取整个body文本,直接定位表格行提取指定行数的单元格内容,匹配你的需求。
额外注意事项
- 若表格选择器不匹配,需打开浏览器开发者工具查看实际HTML结构,替换为对应选择器(如表格的类名、ID)。
- 若仍存在加载延迟,可临时添加
await frame.waitForTimeout(3000)等待3秒,但优先使用waitForSelector确保准确性。
内容的提问来源于stack exchange,提问作者chenzen
相关产品推荐
相关产品推荐

