使用Puppeteer时遭遇‘Execution context was destroyed’错误求助
问题解决:ProtocolError: Execution context was destroyed
问题背景
我用JavaScript结合Puppeteer和Cheerio开发工具,将纽约时报文章URL传入Archive.vn获取完整内容,但运行时持续抛出以下异常:
Exception has occurred: ProtocolError: Protocol error (Runtime.callFunctionOn): Execution context was destroyed.
代码能成功跳转Archive.vn并加载对应页面,但怀疑问题出在代码第36行,更换waitForSelector的目标元素后仍报错。
错误原因分析
这个错误核心是执行page.$eval时,原页面的执行上下文已被销毁,大概率是Archive.vn在加载过程中发生了页面跳转/刷新,导致之前的页面上下文失效。另外代码还有两处明显问题:
- 测试代码里重复声明
articleUrl变量(参数与内部变量重名),会触发语法错误 page.$eval中的选择器错误使用HTML实体",应直接使用正常引号
修正方案
1. 确保页面稳定后再操作
用networkidle2等待页面网络请求稳定,同时更换更准确的Archive.vn存档内容容器作为等待目标,避免等待错误元素导致后续操作提前执行。
2. 修正选择器和变量问题
将选择器中的"替换为正常引号,删除测试代码中重复的变量声明。
修正后的完整代码
import axios from 'axios'; import puppeteer from 'puppeteer'; import cheerio from 'cheerio'; async function getTopStories() { try { const response = await axios.get( 'https://api.nytimes.com/svc/news/v3/content/all/all.json', { params: { 'api-key': 'XXXXXXXX', }, } ); const articleUrl = response.data.results[0].url; console.log('获取到NYT文章URL:', articleUrl); return articleUrl; } catch (error) { console.error('获取NYT头条失败:', error); return null; } } async function archiveArticle(articleUrl) { let browser; try { browser = await puppeteer.launch({ headless: false }); const page = await browser.newPage(); page.setDefaultTimeout(30000); // 导航到Archive.vn并等待页面网络请求稳定 await page.goto(`https://www.archive.vn/?run=1&url=${encodeURIComponent(articleUrl)}`, { waitUntil: 'networkidle2', timeout: 300000 }); console.log('已跳转到archive.vn'); // 等待存档内容容器加载完成 await page.waitForSelector('.archive-body', { timeout: 300000 }); console.log('文章内容已加载'); // 修正选择器引号问题 const articleHtml = await page.$eval('section[name="articleBody"]', (section) => { return section.innerHTML; }); console.log('文章HTML:', articleHtml); const $ = cheerio.load(articleHtml); const articleTitle = $('h1').text(); const articleAuthor = $('.author').text(); const articleDate = $('.date').text(); const articleContent = $('.article-content').text(); console.log(`标题: ${articleTitle}`); console.log(`作者: ${articleAuthor}`); console.log(`日期: ${articleDate}`); console.log(`内容: ${articleContent}`); } catch (error) { console.error('存档文章失败:', error.message); } finally { // 确保浏览器始终关闭 if (browser) await browser.close(); } } (async () => { const nyTimesArticleUrl = await getTopStories(); if (nyTimesArticleUrl) { await archiveArticle(nyTimesArticleUrl); } else { console.error('未能获取NYT文章URL'); } })();
修正后的测试代码
import puppeteer from 'puppeteer'; import cheerio from 'cheerio'; async function archiveArticle() { let browser; try { browser = await puppeteer.launch({ headless: false }); const page = await browser.newPage(); page.setDefaultTimeout(30000); const articleUrl = 'https://www.nytimes.com/2024/04/02/world/middleeast/biden-gaza-aid-attack.html'; await page.goto(`https://www.archive.vn/?run=1&url=${encodeURIComponent(articleUrl)}`, { waitUntil: 'networkidle2', timeout: 300000 }); console.log('已跳转到archive.vn'); await page.waitForSelector('.archive-body', { timeout: 300000 }); console.log('文章内容已加载'); const articleHtml = await page.$eval('section[name="articleBody"]', (section) => { return section.innerHTML; }); console.log('文章HTML:', articleHtml); const $ = cheerio.load(articleHtml); const articleTitle = $('h1').text(); const articleAuthor = $('.author').text(); const articleDate = $('.date').text(); const articleContent = $('.article-content').text(); console.log(`标题: ${articleTitle}`); console.log(`作者: ${articleAuthor}`); console.log(`日期: ${articleDate}`); console.log(`内容: ${articleContent}`); } catch (error) { console.error('存档文章失败:', error.message); } finally { if (browser) await browser.close(); } } archiveArticle();
额外建议
- 若仍遇到上下文销毁问题,可监听页面
framenavigated事件,确认页面不再跳转后再执行后续操作 - 也可先用
page.content()获取整个页面HTML,再用Cheerio解析,彻底避免page.$eval的上下文依赖问题 - 定期检查Archive.vn页面结构,防止选择器因网站更新失效
内容的提问来源于stack exchange,提问作者PandaBlink
相关产品推荐
相关产品推荐

