如何使用Nightmare.js获取带认证的亚马逊链接PDF内容并保存
Hey there! Let's walk through how to properly fetch that password-protected Amazon PDF using NightmareJS. Your current code starts the login flow, but we need to refine it to reliably capture the PDF content after authentication. Here's what you need to do:
Key Improvements & Complete Solution
Instead of relying on a fixed wait(10000) (which can be flaky), we'll wait for specific elements to confirm login success. Then, we'll either use Nightmare's built-in download capability or export cookies to fetch the PDF via an HTTP library (more reliable for binary content).
Option 1: Use Nightmare's download Method
This is the simplest approach if Nightmare handles the authenticated session correctly:
const Nightmare = require('nightmare'); const nightmare = Nightmare({ show: true }); const pdfUrl = 'https://central.amazon.in/fp/reports/lab/17627.pdf?ie=UTF8&&requestID=12297766517607'; nightmare // Navigate to the PDF (will redirect to login) .goto(pdfUrl) // Wait for the email input to load .wait('[name=email]') // Enter credentials .type('[name=email]', 'test@gmail.com') .type('[name=password]', 'test@0') // Click sign-in button .click('#signInSubmit') // Wait for a post-login element (e.g., Amazon's header logo) to confirm success .wait('#nav-logo-sprites') // Download the PDF directly to your local filesystem .download('amazon_report.pdf', pdfUrl) // End the Nightmare session .end() .then(() => console.log('PDF downloaded successfully!')) .catch(err => console.error('Error during process:', err));
Option 2: Export Cookies & Fetch via HTTP Library (More Reliable for Binary Content)
If Nightmare's download method doesn't work as expected, we can extract the authenticated cookies and use request to fetch the PDF:
const Nightmare = require('nightmare'); const fs = require('fs'); const request = require('request'); const pdfUrl = 'https://central.amazon.in/fp/reports/lab/17627.pdf?ie=UTF8&&requestID=12297766517607'; const nightmare = Nightmare({ show: true }); nightmare .goto(pdfUrl) .wait('[name=email]') .type('[name=email]', 'test@gmail.com') .type('[name=password]', 'test@0') .click('#signInSubmit') .wait('#nav-logo-sprites') // Get all cookies from the authenticated session .cookies.get() .end() .then(cookies => { // Convert Nightmare cookies to a request-compatible cookie jar const cookieJar = request.jar(); cookies.forEach(cookie => { cookieJar.setCookie(request.cookie(`${cookie.name}=${cookie.value}`), pdfUrl); }); // Fetch the PDF in binary mode and save to file request({ url: pdfUrl, jar: cookieJar, encoding: null // Critical for preserving binary PDF data }, (err, res, body) => { if (!err && res.statusCode === 200) { fs.writeFileSync('amazon_report.pdf', body); console.log('PDF saved successfully!'); } else { console.error('Failed to fetch PDF:', err || `Status code: ${res.statusCode}`); } }); }) .catch(err => console.error('Nightmare session error:', err));
Important Notes
- Avoid Hardcoded Credentials: Store your email and password in environment variables (e.g.,
process.env.AMAZON_EMAIL) instead of hardcoding them for security. - Handle Anti-Scraping Measures: Amazon may trigger CAPTCHAs or additional login verifications. Since you're using
show: true, you can manually complete these steps when the browser window pops up. - Verify Selectors: Double-check that
#signInSubmitand#nav-logo-spritesare the correct selectors for your Amazon region (they should work for central.amazon.in, but always confirm via browser dev tools).
内容的提问来源于stack exchange,提问作者Parveen yadav

