如何扩展rvest代码处理多层嵌套网页抓取数据生成作者级DataFrame
扩展嵌套层级的网页抓取数据处理方案
问题描述
原方案可处理基础嵌套结构的网页数据抓取,但当前数据存在多层嵌套关系:entry(包含唯一collection)→ book(包含booktitle、year)→ author(包含name,city可能缺失)。需要生成以作者为维度的数据框,每行包含对应collection、booktitle、year、author name、city,最终得到7行数据,同时自动处理缺失字段(如Author 5无city的情况)。
示例网页结构
library(rvest) library(dplyr, warn = FALSE) books <- minimal_html(' <div class="entry"> <div class="collection">Collection 1</div> <div class="book"> <div class="booktitle">Book 1</div> <div class="year">1999</div> <div class="author"> <div class="name">Author 1</div> <div class="city">Austin</div> </div> <div class="author"> <div class="name">Author 2</div> <div class="city">Dallas</div> </div> <div class="author"> <div class="name">Author 3</div> <div class="city">Memphis</div> </div> </div> <div class="book"> <div class="booktitle">Book 2</div> <div class="year">2022</div> <div class="author"> <div class="name">Author 4</div> <div class="city">Houston</div> </div> </div> </div> <div class="entry"> <div class="collection">Collection 2</div> <div class="book"> <div class="booktitle">Book 3</div> <div class="year">1845</div> <div class="author"> <div class="name">Author 5</div> </div> <div class="author"> <div class="name">Author 6</div> <div class="city">Dayton</div> </div> <div class="author"> <div class="name">Author 7</div> <div class="city">Philadelphia</div> </div> </div> </div>')
解决方案代码
data_final <- books %>% html_elements(".entry") %>% lapply(\(entry) { # 获取当前entry对应的collection名称 collection_name <- entry %>% html_element(".collection") %>% html_text2() # 遍历当前entry下的所有书籍 entry %>% html_elements(".book") %>% lapply(\(book) { # 提取单本书的标题和年份 book_title <- book %>% html_element(".booktitle") %>% html_text2() book_year <- book %>% html_element(".year") %>% html_text2() # 遍历单本书下的所有作者,生成作者级别的数据行 book %>% html_elements(".author") %>% lapply(\(author) { tibble( collection = collection_name, booktitle = book_title, year = book_year, author_name = author %>% html_element(".name") %>% html_text2(), author_city = author %>% html_element(".city") %>% html_text2() ) }) %>% bind_rows() }) %>% bind_rows() }) %>% bind_rows()
代码说明
- 从最外层层级遍历:以
entry为起点,确保每个collection能关联到其下所有书籍和作者; - 逐层提取信息:先获取
collection名称,再提取单本书的标题和年份,最后提取每个作者的姓名和城市; - 自动处理缺失值:当元素不存在时,
html_text2()会返回NA,无需额外处理缺失字段; - 逐层合并数据:通过
bind_rows()将作者级别的数据逐步合并,最终得到7行符合要求的完整数据。
内容的提问来源于stack exchange,提问作者bill999
相关产品推荐
相关产品推荐

