You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Puppeteer在元素查找超时后跳过当前步骤并向CSV写入空值

如何让Puppeteer在元素查找超时后跳过当前步骤并向CSV写入空值

嗨,我来帮你搞定这个问题!你的核心需求是让Puppeteer在找不到指定元素时不直接终止脚本,而是写入空值或0,继续处理下一个URL对吧?咱们可以通过**错误捕获(try/catch)**来实现这个目标,同时还要优化下你的CSV写入逻辑,避免重复创建写入器和覆盖文件的问题。

核心思路

page.waitForSelector在找不到元素时会抛出超时错误,只要我们用try/catch块包裹每个元素的查找逻辑,就能捕获这个错误,然后给对应变量赋值null或0,让脚本继续执行下去。另外还要调整CSV的写入方式,避免每次循环都覆盖之前的数据。

修改后的完整代码

const puppeteer = require('puppeteer');
const fs = require('fs');
const csv = require('csv-parser');
const createCsvWriter = require('csv-writer').createObjectCsvWriter;

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  const page = await browser.newPage();
  
  const urls = [];
  // 先读取所有URL到数组,用Promise确保读取完成再继续
  await new Promise((resolve) => {
    fs.createReadStream('urls.csv')
      .pipe(csv())
      .on('data', (row) => {
        urls.push(row.url); // 假设CSV里的列名为'url'
      })
      .on('end', resolve);
  });

  // 初始化数组收集所有爬取记录,避免循环中重复写入覆盖文件
  const allRecords = [];

  // 遍历每个URL
  for (const url of urls) {
    try {
      await page.goto(url, { waitUntil: 'networkidle2', timeout: 10000 }); // 给页面加载也加超时
      const urlVisited = url;
      let element1 = null;
      let element2 = null;

      // 处理第一个元素:用try/catch捕获找不到元素的超时错误
      try {
        const el1 = await page.waitForSelector('xpath/' + 'XPATH', { timeout: 3000 }); // 设置3秒超时,避免久等
        element1 = await page.evaluate(el => el.textContent.trim(), el1);
      } catch (err) {
        console.log(`URL ${url} 找不到第一个元素,赋值为null`);
        element1 = null; // 如果你想写0,直接改成element1 = 0即可
      }

      // 处理第二个元素:同样用try/catch包裹
      try {
        const el2 = await page.waitForSelector('xpath/' + 'XPATH', { timeout: 3000 });
        element2 = await page.evaluate(el => el.textContent.trim(), el2);
      } catch (err) {
        console.log(`URL ${url} 找不到第二个元素,赋值为null`);
        element2 = null;
      }

      // 将当前URL的记录加入总数组
      allRecords.push({
        url: urlVisited,
        price1: element1,
        price2: element2 // 修复了你原代码里key写错的问题(之前写的是price)
      });

    } catch (pageErr) {
      // 处理整个页面加载失败的情况,直接写入空值
      console.log(`URL ${url} 加载失败,跳过`);
      allRecords.push({
        url: url,
        price1: null,
        price2: null
      });
    }
  }

  // 所有URL处理完成后,一次性写入CSV
  const csvWriter = createCsvWriter({
    path: 'output.csv',
    header: [
      { id: 'url', title: 'URL' },
      { id: 'price1', title: 'Price1' },
      { id: 'price2', title: 'Price2' }
    ]
  });

  await csvWriter.writeRecords(allRecords);
  console.log('所有数据已成功写入output.csv');

  await browser.close();
})();

关键改动说明

  • 独立错误捕获:每个元素的查找逻辑都用单独的try/catch包裹,这样某个元素找不到不会影响其他元素的收集,也不会终止整个循环。
  • 合理超时设置:给page.waitForSelector和page.goto都设置了超时时间,既避免无限等待,也不会因为等待太久拖慢效率。
  • CSV写入优化:先把所有爬取记录收集到数组里,等所有URL处理完再一次性写入,避免每次循环都覆盖文件,同时提升写入效率。
  • 细节修复:修正了你原代码里变量名重复(el1用了两次)和CSV记录key错误的问题。

额外提示

如果你想把空值改成0,只需要把代码里的element1 = null改成element1 = 0即可,完全可以根据你的需求调整。

备注:内容来源于stack exchange,提问作者newnewnew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 19:09:28