如何修复Node.js中Puppeteer调用谷歌翻译PDF的超时错误?
问题描述
核心目标:使用Node.js的Puppeteer通过谷歌翻译免费服务翻译单页PDF。
编写的代码如下:
const puppeteer = require('puppeteer'); (async () => { // Launch a headless browser const browser = await puppeteer.launch(); // Open a new page const page = await browser.newPage(); // Replace the URL below with the URL of the PDF you want to translate const pdfUrl = 'https://www.cbu.edu/wp-content/uploads/2020/05/2019-20-ece-computersystems-cs.pdf'; // Navigate to the PDF translation page await page.goto(`https://translate.google.com/translate?hl=en&sl=fr&u=${encodeURIComponent(pdfUrl)}`); // Wait for the translation page to load await page.waitForSelector('#contentframe iframe'); // Get the source URL of the translated content iframe const iframeSrc = await page.evaluate(() => { return document.querySelector('#contentframe iframe').src; }); console.log('Found iframe:', iframeSrc); // Navigate to the translated content iframe await page.goto(iframeSrc); // Wait for translation to complete await page.waitForTimeout(5000); // Adjust the timeout as needed based on the translation complexity // Print the translated content to PDF await page.pdf({ path: '/tmp/translated_pdf.pdf' }); console.log('PDF translation completed successfully.'); // Close the browser await browser.close(); })();
运行代码后出现超时错误:
richardsonoge@richardsonoge-blooglet:/opt/lampp/htdocs/pdf_translator$ node translate_pdf.js /opt/lampp/htdocs/pdf_translator/node_modules/puppeteer-core/lib/cjs/puppeteer/common/WaitTask.js:50 this.#timeoutError = new Errors_js_1.TimeoutError(`Waiting failed: ${options.timeout}ms exceeded`); ^ TimeoutError: Waiting for selector `#contentframe iframe` failed: Waiting failed: 30000ms exceeded at new WaitTask (/opt/lampp/htdocs/pdf_translator/node_modules/puppeteer-core/lib/cjs/puppeteer/common/WaitTask.js:50:34) at IsolatedWorld.waitForFunction (/opt/lampp/htdocs/pdf_translator/node_modules/puppeteer-core/lib/cjs/puppeteer/api/Realm.js:25:26) at PQueryHandler.waitFor (/opt/lampp/htdocs/pdf_translator/node_modules/puppeteer-core/lib/cjs/puppeteer/common/QueryHandler.js:170:95) at runNextTicks (node:internal/process/task_queues:60:5) at process.processImmediate (node:internal/timers:442:9) at async CdpFrame.waitForSelector (/opt/lampp/htdocs/pdf_translator/node_modules/puppeteer-core/lib/cjs/puppeteer/api/Frame.js:468:21) at async CdpPage.waitForSelector (/opt/lampp/htdocs/pdf_translator/node_modules/puppeteer-core/lib/cjs/puppeteer/api/Page.js:1309:20) at async /opt/lampp/htdocs/pdf_translator/translate_pdf.js:17:3 Node.js v18.13.0 richardsonoge@richardsonoge-blooglet:/opt/lampp/htdocs/pdf_translator$
修复方案
超时问题的核心原因是谷歌翻译页面结构已变更,原代码依赖的#contentframe iframe选择器失效,同时无头浏览器默认配置易被反爬机制识别。以下是具体修复步骤:
1. 配置Puppeteer绕过反爬检测
谷歌翻译会识别无头浏览器,需添加模拟真实浏览器的参数:
const browser = await puppeteer.launch({ headless: 'new', // 新版无头模式更接近真实浏览器行为 args: [ '--no-sandbox', '--disable-setuid-sandbox', '--user-agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"' ] });
2. 更新页面选择器和等待逻辑
谷歌翻译当前将内容放在c-wiz容器内的iframe中,替换原选择器:
// 等待翻译容器加载 await page.waitForSelector('c-wiz'); // 获取翻译后的iframe地址 const iframeSrc = await page.evaluate(() => { const iframe = document.querySelector('c-wiz iframe'); return iframe ? iframe.src : null; }); // 容错处理:未找到iframe时直接退出 if (!iframeSrc) { console.error('无法找到翻译内容的iframe'); await browser.close(); return; }
3. 优化页面跳转与内容等待逻辑
放弃固定超时等待,改为等待页面加载完成和内容渲染:
// 跳转到翻译页面时等待网络空闲 await page.goto(`https://translate.google.com/translate?hl=en&sl=en&u=${encodeURIComponent(pdfUrl)}`, { waitUntil: 'networkidle2' }); // 跳转到iframe后等待页面完全加载 await page.goto(iframeSrc, { waitUntil: 'networkidle2' }); // 等待内容容器渲染完成 await page.waitForSelector('body');
完整修复代码
const puppeteer = require('puppeteer'); (async () => { // 启动带反爬规避配置的浏览器 const browser = await puppeteer.launch({ headless: 'new', args: [ '--no-sandbox', '--disable-setuid-sandbox', '--user-agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"' ] }); const page = await browser.newPage(); const pdfUrl = 'https://www.cbu.edu/wp-content/uploads/2020/05/2019-20-ece-computersystems-cs.pdf'; // 注意:sl参数需匹配PDF实际语言,示例PDF为英语,故设为sl=en await page.goto(`https://translate.google.com/translate?hl=en&sl=en&u=${encodeURIComponent(pdfUrl)}`, { waitUntil: 'networkidle2' }); // 等待翻译容器加载 await page.waitForSelector('c-wiz'); // 获取翻译后的iframe地址 const iframeSrc = await page.evaluate(() => { const iframe = document.querySelector('c-wiz iframe'); return iframe ? iframe.src : null; }); if (!iframeSrc) { console.error('无法找到翻译内容的iframe'); await browser.close(); return; } console.log('找到翻译iframe:', iframeSrc); // 跳转到翻译后的页面 await page.goto(iframeSrc, { waitUntil: 'networkidle2' }); // 等待内容渲染完成 await page.waitForSelector('body'); // 生成PDF,指定格式为A4 await page.pdf({ path: '/tmp/translated_pdf.pdf', format: 'A4' }); console.log('PDF翻译完成,已保存到/tmp/translated_pdf.pdf'); await browser.close(); })();
额外注意事项
- 谷歌翻译页面结构可能随时变更,若再次出错需重新检查DOM结构
- 源语言参数
sl必须与PDF实际语言匹配,原代码中sl=fr(法语)不符合示例PDF的英语属性,已修正为sl=en - 避免频繁批量请求,防止触发谷歌访问限制
内容的提问来源于stack exchange,提问作者Richardson
相关产品推荐
相关产品推荐

