You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JavaScript爬虫入门问题:批量爬取报纸文章失败求助

问题排查与修复方案

你的代码存在几个关键问题,导致批量爬取失败,以下是具体问题和修复方法:

1. 语法错误(隐藏的作用域问题)

在page.evaluate的循环内部错误插入了const notas = [];,这行代码破坏了循环结构,属于语法错误,虽然前半部分链接收集可能暂时运行,但会导致后续代码逻辑异常。需要将这行代码移到正确的作用域(比如async函数内部、链接收集完成之后)。

2. 变量作用域错误

循环结束后执行console.log(laCapital),但laCapital是在单个URL处理的try块内部定义的,外部无法访问,会抛出laCapital is not defined的错误,导致后续流程中断。

3. 低效的CSV写入方式

每次处理一个链接就调用writeRecords写入CSV,频繁的文件IO会降低效率,且可能导致文件内容异常。应该先收集所有爬取到的数据,最后一次性写入。

4. 元素获取容错性不足

直接使用page.$eval获取元素,如果目标元素不存在,会抛出错误中断当前链接的处理。需要改用page.$先判断元素是否存在,再获取内容。


修复后的完整代码

const puppeteer = require('puppeteer');
const createCsvWriter = require('csv-writer').createObjectCsvWriter;

const baseUrl = "https://www.lacapital.com.ar/secciones/laciudad.html/";
const numbers = Array.from({length: 2 + 1}, (_, i) => i);
const paginasCiudad = [];

for (const number of numbers) {
  paginasCiudad.push(baseUrl + number);
};
console.log(paginasCiudad);
console.log("posPag");

(async () => {
    const browser = await puppeteer.launch();
    const page = await browser.newPage();
    const urls = [];
    const notas = []; // 移到正确的作用域

    // 收集所有文章链接
    for (let pagina of paginasCiudad) {    
        await page.goto(pagina, { waitUntil: 'domcontentloaded' });
        const enlaces = await page.evaluate(() => {
            const elements = document.querySelectorAll('article.big-entry-box a.cover-link, article.standard-entry-box a.cover-link, article.medium-entry a.cover-link');
            const links = [];

            for (let element of elements) {
                links.push(element.href);
            }
            
            return links;
        });
        urls.push(...enlaces); // 简化数组拼接
    }
    console.log(`共收集到 ${urls.length} 篇文章链接`);

    try {
        // 遍历所有链接爬取内容
        for (let url of urls) {
            try {
                const newData = {};

                await page.goto(url, { waitUntil: 'domcontentloaded' });
        
                // 容错处理:先检查元素是否存在再获取内容
                const fechaEl = await page.$('span.nota-fecha');
                newData.fecha = fechaEl ? await page.evaluate(el => el.textContent, fechaEl) : '';

                const tituloEl = await page.$('h1.nota-title');
                newData.titulo = tituloEl ? await page.evaluate(el => el.textContent, tituloEl) : '';

                const bajadaEl = await page.$('div.nota-bajada');
                newData.bajada = bajadaEl ? await page.evaluate(el => el.textContent, bajadaEl) : '';

                newData.nota = await page.$$eval('div.article-body p', els => els.map(el => el.textContent).join('\n'));

                const imagenEl = await page.$('picture.preview-img > div.extra-holder > img');
                newData.imagen = imagenEl ? await page.evaluate(el => el.getAttribute('src'), imagenEl) : '';

                const seccionEl = await page.$('div.breadcrumbs.flex-container.align-center');
                newData.seccion = seccionEl ? await page.evaluate(el => el.textContent.trim(), seccionEl) : '';

                const tagsEl = await page.$('div.tags-container.flex-container.flex-wrap');
                newData.tags = tagsEl ? await page.evaluate(el => el.textContent.trim(), tagsEl) : '';

                newData.link = url;
                newData.fuente = 'La Capital';
            
                notas.push(newData);
                console.log(`已爬取:${newData.titulo}`);
            } catch (error) {
                console.log(`爬取链接失败 ${url}:`, error.message);
            }
        }

        // 一次性写入所有数据到CSV
        const csvWriter = createCsvWriter({
            path: 'notas.csv',
            header: [
                {id: 'fecha', title: 'Fecha'},
                {id: 'titulo', title: 'Título'},
                {id: 'bajada', title: 'Bajada'},
                {id: 'nota', title: 'Nota'},
                {id: 'imagen', title: 'Imagen'},
                {id: 'seccion', title: 'Sección'},
                {id: 'tags', title: 'Tags'},
                {id: 'link', title: 'Link'},
                {id: 'fuente', title: 'Fuente'}
            ]
        });

        await csvWriter.writeRecords(notas);
        console.log(`所有爬取完成,共保存 ${notas.length} 篇文章到 notas.csv`);
    } catch (error) {
        console.error('整体流程出错:', error);
    } finally {
        await browser.close();
    }
})();

额外优化点

  • 使用urls.push(...enlaces)替代forEach拼接数组,代码更简洁。
  • 将页面等待策略改为domcontentloaded,比load更快,减少等待时间。
  • 对每个元素增加存在性检查,避免单个元素缺失导致爬取中断。
  • 统一收集所有数据后一次性写入CSV,提升效率和稳定性。

内容的提问来源于stack exchange,提问作者Federico Avalos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 01:42:22