You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NodeJS中使用Puppeteer爬取网页PDF链接仅写入一条至文本文件的技术问题求助

Fix: Your PDF Links File Only Saves One Entry

Hey there! I see exactly what's going on here—your console correctly prints every PDF link you scrape, but only the last one sticks around in pdfLinks.txt. The culprit is how you're using fs.writeFile() in your loop.

The Problem

Every time you call await fs.writeFile("pdfLinks.txt", link);, it completely overwrites the existing file with the current link. By the end of your loop, only the final link you processed is left in the file.

Two Simple Solutions

1. Collect All Links First, Then Write Once (Recommended)

This is more efficient (fewer disk operations) and avoids the overwrite issue entirely. We'll filter all PDF links into an array first, then write them all to the file at once:

const puppeteer = require("puppeteer");
const fs = require("fs/promises");
let myNewURL = "https://www.renault.co.il/cars/Zoe/index.html?fbclid=IwAR1RtxbC_U2fImp9_KXJuQ869h5Wv77fyZVj8uBOU86rU90wb2L_NfrNppc";

async function scrapeSite(url) {
 console.log("firing");
 const browser = await puppeteer.launch({ headless: true });
 const page = await browser.newPage();
 await page.goto(url);

 // Grab all links from the page
 const linkCollection = await page.$$eval("a", (links) => {
 return links.map((link) => link.href);
 });

 // Filter to keep only PDF links
 const pdfLinks = linkCollection.filter(link => link.includes(".pdf"));

 // Print all found PDF links (same as before)
 pdfLinks.forEach(link => console.log(link));

 // Write all links to the file, each on a new line
 await fs.writeFile("pdfLinks.txt", pdfLinks.join('\n'));

 await browser.close();
}

scrapeSite(myNewURL);

2. Append Links Instead of Overwriting

If you prefer to keep the loop structure, switch from writeFile to appendFile, which adds content to the end of the file instead of replacing it. You might also want to clear the file first if you run the script multiple times:

const puppeteer = require("puppeteer");
const fs = require("fs/promises");
let myNewURL = "https://www.renault.co.il/cars/Zoe/index.html?fbclid=IwAR1RtxbC_U2fImp9_KXJuQ869h5Wv77fyZVj8uBOU86rU90wb2L_NfrNppc";

async function scrapeSite(url) {
 console.log("firing");
 const browser = await puppeteer.launch({ headless: true });
 const page = await browser.newPage();
 await page.goto(url);

 const linkCollection = await page.$$eval("a", (links) => {
 return links.map((link) => link.href);
 });

 // Clear the file first to avoid old content from previous runs
 await fs.writeFile("pdfLinks.txt", "");

 for (const link of linkCollection) {
 if (link.includes(".pdf")) {
 console.log(link);
 // Append the link plus a newline to the file
 await fs.appendFile("pdfLinks.txt", link + '\n');
 }
 }

 await browser.close();
}

scrapeSite(myNewURL);

Why the First Solution is Better

Writing to disk is a relatively slow operation, so doing it once instead of once per link will make your script run faster, especially if you're scraping a page with lots of PDF links. It also keeps your code cleaner by separating the scraping/filtering logic from the file-writing logic.

Give either approach a try, and you'll see all your PDF links show up in the text file!

内容的提问来源于stack exchange,提问作者Gavin Coulson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 09:57:37