You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Puppeteer+Node.js将关联数据存入MongoDB对应文档而非独立数组

解决Puppeteer抓取时H2标题与关联内容的匹配问题

问题背景

给定如下HTML结构的页面:

<h1>BIG HEADER</h1>
<h2 class="header">a Header</h2>
<p>some content</p>
<br>
<h2 class="header">a Header</h2>
<p>some content</p>
<br>
<h2 class="header">a Header</h2>
<p>some content</p>
<h3 class="small-header">any header</h3>
<p>some more content</p>
<h3 class="small-header">any header</h3>
<p>some more content</p>
<h3 class="small-header">any header</h3>
<p>some more content</p>
<br>

需要将每个H2标题及其关联的内容(包括后续的P、H3+P组合)存入MongoDB,每个H2对应一个独立文档,格式如下:

[
  {
    title: "blogpost-title",
    header: "a header",
    content: "some content"
  },
  {
    title: "blogpost-title",
    header: "a header",
    content: "some content\n\nany header\n\nsome more content\n\nany header\n\nsome more content\n\nany header\n\nsome more content"
  }
]

此前的代码仅能分别抓取H2和P标签文本,无法建立对应关联。

解决方案

核心思路是直接在页面上下文(浏览器端)遍历每个H2元素,通过DOM的兄弟元素关系,收集该H2之后直到下一个H2或BR标签的所有关联内容,从而确保标题与内容的对应关系。

完整代码示例

// 获取页面标题
const pageTitle = await page.title();

// 遍历所有class为header的H2,收集每个H2对应的关联内容
const blogSections = await page.$$eval('h2.header', (h2Elements, pageTitle) => {
  return h2Elements.map(h2 => {
    const headerText = h2.textContent.trim();
    let content = '';
    let nextEl = h2.nextElementSibling;

    // 遍历当前H2的后续兄弟元素,直到遇到下一个H2或BR标签
    while (nextEl) {
      if (nextEl.tagName === 'H2' || nextEl.tagName === 'BR') {
        break;
      }
      // 收集P和H3标签的文本内容
      if (['P', 'H3'].includes(nextEl.tagName)) {
        content += `${nextEl.textContent.trim()}\n\n`;
      }
      nextEl = nextEl.nextElementSibling;
    }

    return {
      title: pageTitle,
      header: headerText,
      content: content.trim() // 移除首尾多余的换行和空格
    };
  });
}, pageTitle); // 将页面标题传入浏览器上下文

// 将结果插入MongoDB(示例代码,需确保已建立MongoDB连接)
// const MongoClient = require('mongodb').MongoClient;
// const client = await MongoClient.connect('mongodb://localhost:27017');
// const db = client.db('your-db-name');
// const collection = db.collection(pageTitle);
// await collection.insertMany(blogSections);
// await client.close();

代码说明

  1. $$eval的使用:通过$$eval直接在浏览器端操作DOM,避免了前后端数据分离导致的关联问题。
  2. 兄弟元素遍历:对每个H2,从其下一个兄弟元素开始遍历,直到遇到下一个H2或BR标签停止,确保只收集当前H2的关联内容。
  3. 内容收集:将遍历到的P和H3标签文本按顺序拼接,保留内容的层级结构。

内容的提问来源于stack exchange,提问作者VebDav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 22:05:21