You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Node.js WriteStream正确写入多页爬取的JSON数组?

Fixing JSON Formatting When Incrementally Writing Scraper Data to a Stream in Node.js

Hey there! I totally get the frustration—when you're scraping pages and writing each array directly to a stream, you end up with invalid JSON because you're just concatenating separate arrays instead of merging their contents. Let's break down two solid solutions depending on your use case.

Option 1: Collect All Data First (Simple, Good for Small-to-Medium Datasets)

If your total dataset isn't massive (won't eat up too much memory), the easiest fix is to collect all your scraped items into a single array first, then write the whole thing to the file once you're done scraping. This avoids any stream-related formatting headaches.

Here's how you'd adjust your code:

const fs = require('fs');
const allScrapedItems = [];

// Your existing page-scraping function (example)
async function scrapeSinglePage(pageNumber) {
  // Replace with your actual scraping logic to get an array of objects
  return [{ id: pageNumber, data: `Page ${pageNumber} content` }];
}

async function runScraper() {
  const totalPages = 5; // Replace with your actual total page count
  
  for (let page = 1; page <= totalPages; page++) {
    const pageItems = await scrapeSinglePage(page);
    // Merge the page's items into the main array (don't push the array itself!)
    allScrapedItems.push(...pageItems);
  }

  // Write the complete array to file
  fs.writeFileSync('scraped-data.json', JSON.stringify(allScrapedItems, null, 2));
  console.log('Scraping done! JSON file is ready.');
}

runScraper();

Option 2: Incremental Stream Writing (Memory-Friendly for Large Datasets)

If you're dealing with a huge amount of data and can't load everything into memory at once, you can manually control the JSON structure as you write to the stream. The key is to write the opening [ first, then write each item individually (with commas between them), and finally close with ].

Check out this implementation:

const fs = require('fs');
const outputStream = fs.createWriteStream('scraped-data.json');
let isFirstItem = true;

// Write the opening bracket of the JSON array
outputStream.write('[');

async function scrapeSinglePage(pageNumber) {
  // Replace with your actual scraping logic
  return [{ id: pageNumber, data: `Page ${pageNumber} content` }];
}

async function runScraper() {
  const totalPages = 5;
  
  try {
    for (let page = 1; page <= totalPages; page++) {
      const pageItems = await scrapeSinglePage(page);
      
      for (const item of pageItems) {
        if (!isFirstItem) {
          // Add a comma before every item except the first one
          outputStream.write(',');
        }
        // Write the individual item as JSON
        outputStream.write(JSON.stringify(item, null, 2));
        isFirstItem = false;
      }
    }

    // Close the JSON array and finish the stream
    outputStream.write(']');
    outputStream.end();
    console.log('Scraping complete! JSON file is valid.');
  } catch (error) {
    console.error('Scraping failed:', error);
    // Make sure to close the stream even if there's an error
    outputStream.end();
  }
}

runScraper();

Key Notes for the Stream Approach:

  • Comma Handling: Always check if you're writing the first item to avoid a leading comma (which would break the JSON).
  • Error Handling: Don't forget to close the stream if an error occurs—otherwise, your file might get stuck in an incomplete state.
  • Asynchronous Flow: Since scraping is usually async, make sure your loop waits for each page to finish before moving on (using await like in the example).

Both approaches will give you a valid single JSON array with all your scraped items merged together—pick the one that fits your dataset size best!

内容的提问来源于stack exchange,提问作者ClementParis016

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:12:35