NodeJS中使用Puppeteer爬取网页PDF链接仅写入一条至文本文件的技术问题求助
Hey there! I see exactly what's going on here—your console correctly prints every PDF link you scrape, but only the last one sticks around in pdfLinks.txt. The culprit is how you're using fs.writeFile() in your loop.
The Problem
Every time you call await fs.writeFile("pdfLinks.txt", link);, it completely overwrites the existing file with the current link. By the end of your loop, only the final link you processed is left in the file.
Two Simple Solutions
1. Collect All Links First, Then Write Once (Recommended)
This is more efficient (fewer disk operations) and avoids the overwrite issue entirely. We'll filter all PDF links into an array first, then write them all to the file at once:
const puppeteer = require("puppeteer"); const fs = require("fs/promises"); let myNewURL = "https://www.renault.co.il/cars/Zoe/index.html?fbclid=IwAR1RtxbC_U2fImp9_KXJuQ869h5Wv77fyZVj8uBOU86rU90wb2L_NfrNppc"; async function scrapeSite(url) { console.log("firing"); const browser = await puppeteer.launch({ headless: true }); const page = await browser.newPage(); await page.goto(url); // Grab all links from the page const linkCollection = await page.$$eval("a", (links) => { return links.map((link) => link.href); }); // Filter to keep only PDF links const pdfLinks = linkCollection.filter(link => link.includes(".pdf")); // Print all found PDF links (same as before) pdfLinks.forEach(link => console.log(link)); // Write all links to the file, each on a new line await fs.writeFile("pdfLinks.txt", pdfLinks.join('\n')); await browser.close(); } scrapeSite(myNewURL);
2. Append Links Instead of Overwriting
If you prefer to keep the loop structure, switch from writeFile to appendFile, which adds content to the end of the file instead of replacing it. You might also want to clear the file first if you run the script multiple times:
const puppeteer = require("puppeteer"); const fs = require("fs/promises"); let myNewURL = "https://www.renault.co.il/cars/Zoe/index.html?fbclid=IwAR1RtxbC_U2fImp9_KXJuQ869h5Wv77fyZVj8uBOU86rU90wb2L_NfrNppc"; async function scrapeSite(url) { console.log("firing"); const browser = await puppeteer.launch({ headless: true }); const page = await browser.newPage(); await page.goto(url); const linkCollection = await page.$$eval("a", (links) => { return links.map((link) => link.href); }); // Clear the file first to avoid old content from previous runs await fs.writeFile("pdfLinks.txt", ""); for (const link of linkCollection) { if (link.includes(".pdf")) { console.log(link); // Append the link plus a newline to the file await fs.appendFile("pdfLinks.txt", link + '\n'); } } await browser.close(); } scrapeSite(myNewURL);
Why the First Solution is Better
Writing to disk is a relatively slow operation, so doing it once instead of once per link will make your script run faster, especially if you're scraping a page with lots of PDF links. It also keeps your code cleaner by separating the scraping/filtering logic from the file-writing logic.
Give either approach a try, and you'll see all your PDF links show up in the text file!
内容的提问来源于stack exchange,提问作者Gavin Coulson

