如何用Puppeteer从Chrome PDF预览界面下载PDF文件
Puppeteer 自动化下载动态生成的PDF解决方案
需求
我需要编写一个Puppeteer脚本完成以下自动化操作:
- 访问指定页面链接
- 点击页面中的「NCAA Tournament Printable Bracket」链接,该链接会在Chrome的PDF预览窗口加载文档
- 保存这个PDF文件
补充说明:该PDF是点击链接后动态生成的,无法直接通过固定URL下载。
已尝试的方法及问题
- 使用
page.pdf()方法:会把当前网页(而非PDF预览窗口)打印成PDF,不符合需求 - 按下
Cmd + A:Mac系统下Chrome的原生功能,Puppeteer无法触发 - Tab键切换到保存按钮再按Enter:Mac的系统保存弹窗对Puppeteer不可见,操作无效
尝试过的代码
const puppeteer = require('puppeteer'); async function run() { // const browser = await puppeteer.launch(); const browser = await puppeteer.launch({ headless: false }); const page = await browser.newPage(); await page.goto('https://www.espn.com/sports-betting/story/_/id/35852378/2023-ncaa-men-tournament-march-madness-printable-brackets-college-basketball'); const html = await page.content(); await page.waitForSelector('#article-feed > article:nth-child(1) > div > div.article-body > ul > li:nth-child(1) > p > a > em > strong'); // Print and save test 1 await page.pdf({ path: 'bracket.pdf', format: 'A4' }); // Print and save test 2 await page.keyboard.press('Tab'); await page.keyboard.press('Tab'); await page.keyboard.press('Tab'); await page.keyboard.press('Tab'); await page.keyboard.press('Tab'); await page.keyboard.press('Tab'); await page.keyboard.press('Tab'); await page.keyboard.press('Tab'); await page.keyboard.press('Enter'); await page.keyboard.press('Enter'); // Print and save test 3 await page.evaluate(() => { window.print = function() {}; }); // Print and save test 4 await page.keyboard.down('Cmd'); await page.keyboard.down('S'); await page.keyboard.up('S'); await page.keyboard.up('Cmd'); await page.waitForTimeout(5000); await page.keyboard.up('Enter'); await browser.close(); } run();
解决方案
核心思路是拦截点击链接后的请求,直接获取PDF的真实地址,绕开系统弹窗交互:
const puppeteer = require('puppeteer'); const fs = require('fs'); async function run() { const browser = await puppeteer.launch({ headless: false }); const page = await browser.newPage(); // 开启请求拦截,监听PDF格式请求 await page.setRequestInterception(true); let pdfUrl = null; page.on('request', (request) => { // 筛选PDF请求,记录地址后拦截默认预览行为 if (request.url().endsWith('.pdf')) { pdfUrl = request.url(); request.abort(); } else { request.continue(); } }); // 访问目标页面,等待网络状态稳定 await page.goto('https://www.espn.com/sports-betting/story/_/id/35852378/2023-ncaa-men-tournament-march-madness-printable-brackets-college-basketball', { waitUntil: 'networkidle2' }); // 等待目标链接加载完成并点击 const bracketLink = await page.waitForSelector('#article-feed > article:nth-child(1) > div > div.article-body > ul > li:nth-child(1) > p > a'); await bracketLink.click(); // 等待获取到PDF地址 while (!pdfUrl) { await page.waitForTimeout(100); } // 下载并保存PDF文件 const pdfBuffer = await page.goto(pdfUrl); await fs.promises.writeFile('ncaa-bracket.pdf', await pdfBuffer.buffer()); await browser.close(); } run();
方案说明
- 请求拦截:通过
setRequestInterception监听所有请求,直接捕获PDF的真实下载地址 - 绕开系统弹窗:跳过Chrome的PDF预览环节,直接通过获取到的URL下载文件,避免和不可见的系统弹窗交互
- 稳定操作:使用
waitForSelector和networkidle2确保页面元素加载完成后再执行点击,避免操作时机错误
内容的提问来源于stack exchange,提问作者SuperTony
相关产品推荐
相关产品推荐

