使用rvest爬取Quotes To Scrape时无法存储作者列表求助
问题:爬取Quotes To Scrape时作者列表无法填充
我尝试将Quotes To Scrape网站的作者信息存储到列表中,爬取名言的操作能正常执行并得到完整的名言列表,但作者列表始终为空。我原本以为authors是全局变量,应该能被填充,但实际并非如此。
我的函数脚本
library(rvest) library(tidyverse) start_url <- 'http://quotes.toscrape.com/' start_page <- read_html(start_url) session <- session(start_url) get_quotes_elements <- function(page_url) { page <- read_html(page_url) quotes_elements <- html_nodes(page,".quote") return(quotes_elements) } get_quote <- function(quote_element) { quote <- list() quote_text = html_nodes(quote_element,'.text') %>% html_text() quote_author = html_nodes(quote_element,'.author') %>% html_text() quote_tags = html_nodes(quote_element,'.tags') %>% html_text() quote['Author'] <- quote_author quote['Quote'] <- quote_text quote['Tags'] <- quote_tags author_page_url <- paste0('https://quotes.toscrape.com', html_nodes(quote_element,"a[href*='/author/']") %>% html_attr("href")) author <- get_author(author_page_url) authors <- append(authors, author) print(paste(" Quote from ", author["Name"], " added")) return(quote) } get_quotes_pages <- function(start_url) { start_page_number <- 1 pages_count <- 9 quotes_pages <- list() quotes_pages <- append(quotes_pages, start_url) for (i in start_page_number:pages_count) { new_quotes_page_url <- quotes_pages[[i]] # print(paste("Processing ", new_quotes_page_url)) new_quotes_page <- read_html(session %>% session_jump_to(new_quotes_page_url)) next_quotes_page_url <- paste0('https://quotes.toscrape.com', new_quotes_page %>% html_nodes('li.next a') %>% html_attr("href")) # print(paste("next_quotes_page_url = ", next_quotes_page_url)) quotes_pages <- append(quotes_pages, next_quotes_page_url) } return(quotes_pages) } get_author <- function(author_page_url) { author <- list() first_author_page <- read_html(session %>% session_jump_to(author_page_url)) author_name = html_nodes(first_author_page,'.author-title') %>% html_text() # print(paste(" Author name ", author_name)) author_born_date = html_nodes(first_author_page,'.author-born-date') %>% html_text() # print(paste(" Author born date ", author_born_date)) author_description = html_nodes(first_author_page,'.author-description') %>% html_text() # print(paste(" Author description ", author_description)) author['Name'] <- author_name author['BornDate'] <- author_born_date author['Description'] <- author_description return(author) }
运行脚本(示例限制为第一页名言)
quotes_pages <- get_quotes_pages(start_url) first_quotes_page <- quotes_pages[[1]] quotes <- list() authors <- list() for (quotes_page in first_quotes_page) { print(paste("Processing ", quotes_page)) new_quotes <- list() new_quotes <- lapply(get_quotes_elements(quotes_page), get_quote) quotes <- append(quotes, new_quotes) }
问题原因
在get_quote函数中,你直接对authors执行append操作,但R默认将函数内的变量视为局部变量——即便全局环境存在同名的authors,函数内的authors会被当成新的局部变量创建,不会修改全局的列表。
解决方法
方法1:使用全局变量赋值符
在get_quote函数里,把authors <- append(authors, author)改成:
authors <<- append(authors, author)
<<-赋值符会直接修改全局环境中的authors变量,而非创建局部变量。
方法2:返回作者信息并外部收集(更推荐)
避免依赖全局变量,让get_quote同时返回名言和作者信息,再在外部循环中分别收集:
- 修改
get_quote函数:
get_quote <- function(quote_element) { quote <- list() quote_text = html_nodes(quote_element,'.text') %>% html_text() quote_author = html_nodes(quote_element,'.author') %>% html_text() quote_tags = html_nodes(quote_element,'.tags') %>% html_text() quote['Author'] <- quote_author quote['Quote'] <- quote_text quote['Tags'] <- quote_tags author_page_url <- paste0('https://quotes.toscrape.com', html_nodes(quote_element,"a[href*='/author/']") %>% html_attr("href")) author <- get_author(author_page_url) print(paste(" Quote from ", author["Name"], " added")) # 返回包含名言和作者的列表 return(list(quote = quote, author = author)) }
- 修改运行脚本:
quotes_pages <- get_quotes_pages(start_url) first_quotes_page <- quotes_pages[[1]] quotes <- list() authors <- list() for (quotes_page in first_quotes_page) { print(paste("Processing ", quotes_page)) # 获取包含名言和作者的结果集 results <- lapply(get_quotes_elements(quotes_page), get_quote) # 分别提取并追加到对应列表 quotes <- append(quotes, lapply(results, function(x) x$quote)) authors <- append(authors, lapply(results, function(x) x$author)) }
这种方式更符合函数式编程规范,代码模块化更强,也更易调试维护。
额外优化建议
get_quotes_pages中硬编码pages_count <- 9不够灵活,建议通过判断页面是否存在「下一页」按钮来循环获取所有页面,避免固定页数导致的遗漏或错误。- 已经全局创建了
session对象,函数内可以直接用session调用节点,无需每次session_jump_to后再read_html,简化代码。
内容的提问来源于stack exchange,提问作者Didier mac cormick
相关产品推荐
相关产品推荐

