You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Puppeteer连接Bright Data时遭遇Page.navigate limit reached错误求助

解决Puppeteer连接Bright Data时的"Page.navigate limit reached"错误

问题背景

我编写了Puppeteer脚本,用于从存储医生姓名及主页链接的JSON文件读取数据,遍历访问医生主页以提取其官网信息。但使用Bright Data代理连接时,无论尝试单页调用、递归循环、添加导航延迟等方法,始终返回**"Page.navigate limit reached"**错误;而无代理连接的版本可正常运行,且Bright Data日志显示无异常,现寻求该错误的解决方法。

医生数据JSON示例

[
  {
    "name": "Herr Prof. Dr. med. Armin Quentmeyer",
    "link": "https://www.doctolib.de/orthopadie/ludwigshafen-am-rhein/dr-prof-armin-quentmeyer?pid=practice-238444"
  },
  {
    "name": "Herr Dr. med. Ralph Hower",
    "link": "https://www.doctolib.de/orthopadie/murnau-am-staffelsee/ralph-hower?pid=practice-228573"
  },
  {
    "name": "Orthopädische Praxis Ludwigsburg",
    "link": "https://www.doctolib.de/praxis/ludwigsburg/orthopaedische-praxis-ludwigsburg?pid=practice-570039"
  }
]

注:原示例JSON存在语法错误,已修正为数组格式

连接Bright Data的问题脚本

// website_adder.js
import puppeteer from "puppeteer-core";
import fs from 'fs';

async function addWebsites() {
    let browser;
    try {
        console.log("Starting function");
        const auth = '***';
        
        browser = await puppeteer.connect({
            browserWSEndpoint: `wss://${auth}@***`
        });

        // Read the JSON file with the list of doctors
        const data = fs.readFileSync('uniqueDoctorsData.json', 'utf-8');
        const doctorsData = JSON.parse(data);

        // Iterate through each doctor, visit their profile, and add the website to their data
        for (let doctor of doctorsData) {
            console.log(`Visiting profile of ${doctor.name} at ${doctor.link}`);

            try {
                // Introduce a short delay between navigations to avoid hitting the page limit
                await new Promise(resolve => setTimeout(resolve, 2000)); // Wait for 10 seconds
                console.log("Waited for 2000ms");

                // Use a new page for each doctor's profile
                const profilePage = await browser.newPage();
                await profilePage.goto(doctor.link, { waitUntil: 'networkidle2' });

                // Extract the doctor's website from their profile page
                const website = await profilePage.evaluate(() => {
                    const websiteElement = document.querySelector('a[rel="nofollow"][target="_blank"]');
                    return websiteElement ? websiteElement.href : null;
                });

                // Add the website to the doctor's data
                doctor.website = website;

                console.log(`Extracted website for ${doctor.name}: ${website}`);

                // Close the profile page to free up resources
                await profilePage.close();

                // Save the updated doctors data back to the JSON file immediately after processing each doctor
                fs.writeFileSync('uniqueDoctorsData.json', JSON.stringify(doctorsData, null, 2), 'utf-8');
                console.log(`Saved updated data for ${doctor.name} to JSON file`);

            } catch (error) {
                console.error(`Failed to navigate to ${doctor.link}`, error);
            }
        }

        console.log('All doctor profiles processed and data saved to updatedDoctorsData.json');

    } catch (e) {
        console.error('Failed to add websites', e);
    } finally {
        if (browser) {
            await browser.close();
        }
    }
}

// Export the function to be used in index.js
export { addWebsites };

无代理正常运行的脚本

// website_adder.js
import puppeteer from "puppeteer";
import fs from 'fs';

async function addWebsites() {
    let browser;
    try {
        console.log("Starting function");

        // Launch a new browser instance (no proxy connection)
        browser = await puppeteer.launch({
            headless: true, // Run in headless mode for efficiency
            args: ['--no-sandbox', '--disable-setuid-sandbox'] // Additional arguments to improve performance/stability
        });

        // Read the JSON file with the list of doctors
        const data = fs.readFileSync('uniqueDoctorsData.json', 'utf-8');
        const doctorsData = JSON.parse(data);

        // Iterate through each doctor, visit their profile, and add the website to their data
        for (let doctor of doctorsData) {
            // Skip doctors who already have a website to avoid unnecessary work
            if (doctor.website) {
                console.log(`Skipping ${doctor.name} as website is already added.`);
                continue;
            }

            console.log(`Visiting profile of ${doctor.name} at ${doctor.link}`);

            try {
                // Introduce a short delay between navigations to avoid hitting the page limit
                await new Promise(resolve => setTimeout(resolve, 2000)); // Wait for 2 seconds
                console.log("Waited for 2000ms");

                // Use a new page for each doctor's profile
                const profilePage = await browser.newPage();
                await profilePage.goto(doctor.link, { waitUntil: 'networkidle2' });

                // Extract the doctor's website from their profile page
                const website = await profilePage.evaluate(() => {
                    const websiteElement = document.querySelector('a[rel="nofollow"][target="_blank"]');
                    return websiteElement ? websiteElement.href : null;
                });

                // Add the website to the doctor's data
                doctor.website = website;

                console.log(`Extracted website for ${doctor.name}: ${website}`);

                // Close the profile page to free up resources
                await profilePage.close();

                // Save the updated doctors data back to the JSON file immediately after processing each doctor
                fs.writeFileSync('uniqueDoctorsData.json', JSON.stringify(doctorsData, null, 2), 'utf-8');
                console.log(`Saved updated data for ${doctor.name} to JSON file`);

            } catch (error) {
                console.error(`Failed to navigate to ${doctor.link}`, error);
            }
        }

        console.log('All doctor profiles processed and data saved to uniqueDoctorsData.json');

    } catch (e) {
        console.error('Failed to add websites', e);
    } finally {
        if (browser) {
            await browser.close();
        }

        // After processing the current batch, check again if there are any entries left without websites
        const updatedData = fs.readFileSync('uniqueDoctorsData.json', 'utf-8');
        const updatedDoctorsData = JSON.parse(updatedData);
        const remainingDoctors = updatedDoctorsData.filter(doctor => !doctor.website);

        if (remainingDoctors.length > 0) {
            console.log("Continuing to process remaining doctors without websites...");
            // Call the function again to process remaining entries
            await addWebsites();
        } else {
            console.log("All doctors have websites. Process completed.");
        }
    }
}

// Export the function to be used in index.js
export { addWebsites };

解决方法

1. 复用单个页面而非每次新建

Bright Data的会话对页面导航次数或页面实例数有限制,每次循环新建page会快速耗尽配额。修改为复用单个页面:

// 修改Bright Data版本脚本的循环部分
async function addWebsites() {
    let browser;
    let profilePage; // 声明页面变量在循环外
    try {
        // ... 连接浏览器、读取数据等代码不变

        // 提前创建单个页面
        profilePage = await browser.newPage();

        for (let doctor of doctorsData) {
            if (doctor.website) {
                console.log(`Skipping ${doctor.name} as website is already added.`);
                continue;
            }
            console.log(`Visiting profile of ${doctor.name} at ${doctor.link}`);
            try {
                // 延长并随机化延迟,避免固定间隔触发限制
                const delay = Math.floor(Math.random() * 4000) + 3000; // 3-7秒随机延迟
                await new Promise(resolve => setTimeout(resolve, delay));
                console.log(`Waited for ${delay}ms`);

                // 复用页面导航,添加超时设置
                await profilePage.goto(doctor.link, { 
                    waitUntil: 'networkidle2',
                    timeout: 30000 // 30秒超时
                });

                // 提取网站代码不变
                const website = await profilePage.evaluate(() => {
                    const websiteElement = document.querySelector('a[rel="nofollow"][target="_blank"]');
                    return websiteElement ? websiteElement.href : null;
                });

                doctor.website = website;
                console.log(`Extracted website for ${doctor.name}: ${website}`);

                fs.writeFileSync('uniqueDoctorsData.json', JSON.stringify(doctorsData, null, 2), 'utf-8');
                console.log(`Saved updated data for ${doctor.name} to JSON file`);

            } catch (error) {
                console.error(`Failed to navigate to ${doctor.link}`, error);
            }
        }

        // 循环结束后再关闭页面
        await profilePage.close();
        // ... 后续代码不变
    } catch (e) {
        console.error('Failed to add websites', e);
    } finally {
        if (profilePage) await profilePage.close();
        if (browser) await browser.close();
    }
}

2. 调整Bright Data会话配置

  • 检查当前Bright Data套餐的会话限制:部分套餐会限制单个会话内的页面导航次数,可登录后台查看配额或联系客服调整
  • 切换到Unmetered套餐:该套餐无导航次数限制,适合批量爬取场景
  • 调整会话超时时间:延长会话超时,避免频繁重建连接

3. 分批处理数据

将医生数据分成小批次,每处理完一批就关闭浏览器并重新建立连接,避免单个会话内导航次数过多:

// 示例:每5个医生为一批
const batchSize = 5;
for (let i = 0; i < doctorsData.length; i += batchSize) {
    const batch = doctorsData.slice(i, i + batchSize);
    // 新建浏览器连接处理当前批次
    let browser = await puppeteer.connect({ browserWSEndpoint: `wss://${auth}@***` });
    let page = await browser.newPage();
    for (let doctor of batch) {
        // 处理单个医生逻辑
    }
    await page.close();
    await browser.close();
    // 批次间添加更长延迟
    await new Promise(resolve => setTimeout(resolve, 10000));
}

4. 使用Bright Data官方SDK

改用Bright Data提供的专用SDK,它会自动处理会话限制、资源管理和重试逻辑,比直接用Puppeteer连接WSS更稳定。


内容的提问来源于stack exchange,提问作者Galla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 15:24:58