Puppeteer连接Bright Data时遭遇Page.navigate limit reached错误求助
问题背景
我编写了Puppeteer脚本,用于从存储医生姓名及主页链接的JSON文件读取数据,遍历访问医生主页以提取其官网信息。但使用Bright Data代理连接时,无论尝试单页调用、递归循环、添加导航延迟等方法,始终返回**"Page.navigate limit reached"**错误;而无代理连接的版本可正常运行,且Bright Data日志显示无异常,现寻求该错误的解决方法。
医生数据JSON示例
[ { "name": "Herr Prof. Dr. med. Armin Quentmeyer", "link": "https://www.doctolib.de/orthopadie/ludwigshafen-am-rhein/dr-prof-armin-quentmeyer?pid=practice-238444" }, { "name": "Herr Dr. med. Ralph Hower", "link": "https://www.doctolib.de/orthopadie/murnau-am-staffelsee/ralph-hower?pid=practice-228573" }, { "name": "Orthopädische Praxis Ludwigsburg", "link": "https://www.doctolib.de/praxis/ludwigsburg/orthopaedische-praxis-ludwigsburg?pid=practice-570039" } ]
注:原示例JSON存在语法错误,已修正为数组格式
连接Bright Data的问题脚本
// website_adder.js import puppeteer from "puppeteer-core"; import fs from 'fs'; async function addWebsites() { let browser; try { console.log("Starting function"); const auth = '***'; browser = await puppeteer.connect({ browserWSEndpoint: `wss://${auth}@***` }); // Read the JSON file with the list of doctors const data = fs.readFileSync('uniqueDoctorsData.json', 'utf-8'); const doctorsData = JSON.parse(data); // Iterate through each doctor, visit their profile, and add the website to their data for (let doctor of doctorsData) { console.log(`Visiting profile of ${doctor.name} at ${doctor.link}`); try { // Introduce a short delay between navigations to avoid hitting the page limit await new Promise(resolve => setTimeout(resolve, 2000)); // Wait for 10 seconds console.log("Waited for 2000ms"); // Use a new page for each doctor's profile const profilePage = await browser.newPage(); await profilePage.goto(doctor.link, { waitUntil: 'networkidle2' }); // Extract the doctor's website from their profile page const website = await profilePage.evaluate(() => { const websiteElement = document.querySelector('a[rel="nofollow"][target="_blank"]'); return websiteElement ? websiteElement.href : null; }); // Add the website to the doctor's data doctor.website = website; console.log(`Extracted website for ${doctor.name}: ${website}`); // Close the profile page to free up resources await profilePage.close(); // Save the updated doctors data back to the JSON file immediately after processing each doctor fs.writeFileSync('uniqueDoctorsData.json', JSON.stringify(doctorsData, null, 2), 'utf-8'); console.log(`Saved updated data for ${doctor.name} to JSON file`); } catch (error) { console.error(`Failed to navigate to ${doctor.link}`, error); } } console.log('All doctor profiles processed and data saved to updatedDoctorsData.json'); } catch (e) { console.error('Failed to add websites', e); } finally { if (browser) { await browser.close(); } } } // Export the function to be used in index.js export { addWebsites };
无代理正常运行的脚本
// website_adder.js import puppeteer from "puppeteer"; import fs from 'fs'; async function addWebsites() { let browser; try { console.log("Starting function"); // Launch a new browser instance (no proxy connection) browser = await puppeteer.launch({ headless: true, // Run in headless mode for efficiency args: ['--no-sandbox', '--disable-setuid-sandbox'] // Additional arguments to improve performance/stability }); // Read the JSON file with the list of doctors const data = fs.readFileSync('uniqueDoctorsData.json', 'utf-8'); const doctorsData = JSON.parse(data); // Iterate through each doctor, visit their profile, and add the website to their data for (let doctor of doctorsData) { // Skip doctors who already have a website to avoid unnecessary work if (doctor.website) { console.log(`Skipping ${doctor.name} as website is already added.`); continue; } console.log(`Visiting profile of ${doctor.name} at ${doctor.link}`); try { // Introduce a short delay between navigations to avoid hitting the page limit await new Promise(resolve => setTimeout(resolve, 2000)); // Wait for 2 seconds console.log("Waited for 2000ms"); // Use a new page for each doctor's profile const profilePage = await browser.newPage(); await profilePage.goto(doctor.link, { waitUntil: 'networkidle2' }); // Extract the doctor's website from their profile page const website = await profilePage.evaluate(() => { const websiteElement = document.querySelector('a[rel="nofollow"][target="_blank"]'); return websiteElement ? websiteElement.href : null; }); // Add the website to the doctor's data doctor.website = website; console.log(`Extracted website for ${doctor.name}: ${website}`); // Close the profile page to free up resources await profilePage.close(); // Save the updated doctors data back to the JSON file immediately after processing each doctor fs.writeFileSync('uniqueDoctorsData.json', JSON.stringify(doctorsData, null, 2), 'utf-8'); console.log(`Saved updated data for ${doctor.name} to JSON file`); } catch (error) { console.error(`Failed to navigate to ${doctor.link}`, error); } } console.log('All doctor profiles processed and data saved to uniqueDoctorsData.json'); } catch (e) { console.error('Failed to add websites', e); } finally { if (browser) { await browser.close(); } // After processing the current batch, check again if there are any entries left without websites const updatedData = fs.readFileSync('uniqueDoctorsData.json', 'utf-8'); const updatedDoctorsData = JSON.parse(updatedData); const remainingDoctors = updatedDoctorsData.filter(doctor => !doctor.website); if (remainingDoctors.length > 0) { console.log("Continuing to process remaining doctors without websites..."); // Call the function again to process remaining entries await addWebsites(); } else { console.log("All doctors have websites. Process completed."); } } } // Export the function to be used in index.js export { addWebsites };
解决方法
1. 复用单个页面而非每次新建
Bright Data的会话对页面导航次数或页面实例数有限制,每次循环新建page会快速耗尽配额。修改为复用单个页面:
// 修改Bright Data版本脚本的循环部分 async function addWebsites() { let browser; let profilePage; // 声明页面变量在循环外 try { // ... 连接浏览器、读取数据等代码不变 // 提前创建单个页面 profilePage = await browser.newPage(); for (let doctor of doctorsData) { if (doctor.website) { console.log(`Skipping ${doctor.name} as website is already added.`); continue; } console.log(`Visiting profile of ${doctor.name} at ${doctor.link}`); try { // 延长并随机化延迟,避免固定间隔触发限制 const delay = Math.floor(Math.random() * 4000) + 3000; // 3-7秒随机延迟 await new Promise(resolve => setTimeout(resolve, delay)); console.log(`Waited for ${delay}ms`); // 复用页面导航,添加超时设置 await profilePage.goto(doctor.link, { waitUntil: 'networkidle2', timeout: 30000 // 30秒超时 }); // 提取网站代码不变 const website = await profilePage.evaluate(() => { const websiteElement = document.querySelector('a[rel="nofollow"][target="_blank"]'); return websiteElement ? websiteElement.href : null; }); doctor.website = website; console.log(`Extracted website for ${doctor.name}: ${website}`); fs.writeFileSync('uniqueDoctorsData.json', JSON.stringify(doctorsData, null, 2), 'utf-8'); console.log(`Saved updated data for ${doctor.name} to JSON file`); } catch (error) { console.error(`Failed to navigate to ${doctor.link}`, error); } } // 循环结束后再关闭页面 await profilePage.close(); // ... 后续代码不变 } catch (e) { console.error('Failed to add websites', e); } finally { if (profilePage) await profilePage.close(); if (browser) await browser.close(); } }
2. 调整Bright Data会话配置
- 检查当前Bright Data套餐的会话限制:部分套餐会限制单个会话内的页面导航次数,可登录后台查看配额或联系客服调整
- 切换到Unmetered套餐:该套餐无导航次数限制,适合批量爬取场景
- 调整会话超时时间:延长会话超时,避免频繁重建连接
3. 分批处理数据
将医生数据分成小批次,每处理完一批就关闭浏览器并重新建立连接,避免单个会话内导航次数过多:
// 示例:每5个医生为一批 const batchSize = 5; for (let i = 0; i < doctorsData.length; i += batchSize) { const batch = doctorsData.slice(i, i + batchSize); // 新建浏览器连接处理当前批次 let browser = await puppeteer.connect({ browserWSEndpoint: `wss://${auth}@***` }); let page = await browser.newPage(); for (let doctor of batch) { // 处理单个医生逻辑 } await page.close(); await browser.close(); // 批次间添加更长延迟 await new Promise(resolve => setTimeout(resolve, 10000)); }
4. 使用Bright Data官方SDK
改用Bright Data提供的专用SDK,它会自动处理会话限制、资源管理和重试逻辑,比直接用Puppeteer连接WSS更稳定。
内容的提问来源于stack exchange,提问作者Galla
相关产品推荐
相关产品推荐

