You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何扩展rvest代码处理多层嵌套网页抓取数据生成作者级DataFrame

扩展嵌套层级的网页抓取数据处理方案

问题描述

原方案可处理基础嵌套结构的网页数据抓取,但当前数据存在多层嵌套关系:entry(包含唯一collection)→ book(包含booktitle、year)→ author(包含name,city可能缺失)。需要生成以作者为维度的数据框,每行包含对应collection、booktitle、year、author name、city,最终得到7行数据,同时自动处理缺失字段(如Author 5无city的情况)。

示例网页结构

library(rvest)
library(dplyr, warn = FALSE)
books <- minimal_html('
  <div class="entry">
        <div class="collection">Collection 1</div>
        <div class="book">
          <div class="booktitle">Book 1</div>
          <div class="year">1999</div>
          <div class="author">
            <div class="name">Author 1</div>
            <div class="city">Austin</div>
          </div>  
          <div class="author">
            <div class="name">Author 2</div>
            <div class="city">Dallas</div> 
          </div>  
          <div class="author">
            <div class="name">Author 3</div>
            <div class="city">Memphis</div>  
          </div>  
        </div>
        <div class="book">
          <div class="booktitle">Book 2</div>
          <div class="year">2022</div>
          <div class="author">
            <div class="name">Author 4</div>
            <div class="city">Houston</div>
          </div>  
        </div>
  </div>
  <div class="entry">  
        <div class="collection">Collection 2</div>
        <div class="book">
          <div class="booktitle">Book 3</div>
          <div class="year">1845</div>
          <div class="author">
            <div class="name">Author 5</div> 
          </div>  
          <div class="author">
            <div class="name">Author 6</div>
            <div class="city">Dayton</div>
          </div>  
          <div class="author">
            <div class="name">Author 7</div>
            <div class="city">Philadelphia</div>  
          </div>  
        </div>
  </div>')

解决方案代码

data_final <- books %>%
  html_elements(".entry") %>%
  lapply(\(entry) {
    # 获取当前entry对应的collection名称
    collection_name <- entry %>% html_element(".collection") %>% html_text2()
    
    # 遍历当前entry下的所有书籍
    entry %>% html_elements(".book") %>%
      lapply(\(book) {
        # 提取单本书的标题和年份
        book_title <- book %>% html_element(".booktitle") %>% html_text2()
        book_year <- book %>% html_element(".year") %>% html_text2()
        
        # 遍历单本书下的所有作者,生成作者级别的数据行
        book %>% html_elements(".author") %>%
          lapply(\(author) {
            tibble(
              collection = collection_name,
              booktitle = book_title,
              year = book_year,
              author_name = author %>% html_element(".name") %>% html_text2(),
              author_city = author %>% html_element(".city") %>% html_text2()
            )
          }) %>% bind_rows()
      }) %>% bind_rows()
  }) %>% bind_rows()

代码说明

  1. 从最外层层级遍历:以entry为起点,确保每个collection能关联到其下所有书籍和作者;
  2. 逐层提取信息:先获取collection名称,再提取单本书的标题和年份,最后提取每个作者的姓名和城市;
  3. 自动处理缺失值:当元素不存在时,html_text2()会返回NA,无需额外处理缺失字段;
  4. 逐层合并数据:通过bind_rows()将作者级别的数据逐步合并,最终得到7行符合要求的完整数据。

内容的提问来源于stack exchange,提问作者bill999

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 22:43:26