You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest爬取Quotes To Scrape时无法存储作者列表求助

问题:爬取Quotes To Scrape时作者列表无法填充

我尝试将Quotes To Scrape网站的作者信息存储到列表中,爬取名言的操作能正常执行并得到完整的名言列表,但作者列表始终为空。我原本以为authors是全局变量,应该能被填充,但实际并非如此。

我的函数脚本

library(rvest)
library(tidyverse)

start_url <- 'http://quotes.toscrape.com/'
start_page <- read_html(start_url)
session <- session(start_url)

get_quotes_elements <- function(page_url) {
  page <- read_html(page_url)
  quotes_elements <- html_nodes(page,".quote")
  return(quotes_elements)
}

get_quote <- function(quote_element) {
  quote <- list()
  quote_text = html_nodes(quote_element,'.text') %>% html_text()
  quote_author = html_nodes(quote_element,'.author') %>% html_text()
  quote_tags = html_nodes(quote_element,'.tags') %>% html_text()
  quote['Author'] <- quote_author
  quote['Quote'] <- quote_text
  quote['Tags'] <- quote_tags
  author_page_url <- paste0('https://quotes.toscrape.com', html_nodes(quote_element,"a[href*='/author/']") %>% html_attr("href"))
  author <- get_author(author_page_url)
  authors <- append(authors, author)
  print(paste(" Quote from ", author["Name"], " added"))
  return(quote)
}

get_quotes_pages <- function(start_url) {
  start_page_number <- 1
  pages_count <- 9
  quotes_pages <- list()
  quotes_pages <- append(quotes_pages, start_url)
  for (i in start_page_number:pages_count) {
    new_quotes_page_url <- quotes_pages[[i]]
    # print(paste("Processing ", new_quotes_page_url))
    new_quotes_page <- read_html(session %>% session_jump_to(new_quotes_page_url))
    next_quotes_page_url <- paste0('https://quotes.toscrape.com', new_quotes_page %>% html_nodes('li.next a') %>% html_attr("href"))
    # print(paste("next_quotes_page_url =  ", next_quotes_page_url))
    quotes_pages <- append(quotes_pages, next_quotes_page_url)
  }
  return(quotes_pages)
}

get_author <- function(author_page_url) {
  author <- list()
  first_author_page <- read_html(session %>% session_jump_to(author_page_url))
  author_name = html_nodes(first_author_page,'.author-title') %>% html_text()
  # print(paste("  Author name  ", author_name))
  author_born_date = html_nodes(first_author_page,'.author-born-date') %>% html_text()
  # print(paste("  Author born date  ", author_born_date))
  author_description = html_nodes(first_author_page,'.author-description') %>% html_text()
  # print(paste("  Author description  ", author_description))
  author['Name'] <- author_name
  author['BornDate'] <- author_born_date
  author['Description'] <- author_description
  return(author)
}

运行脚本(示例限制为第一页名言)

quotes_pages <- get_quotes_pages(start_url)
first_quotes_page <- quotes_pages[[1]]
quotes <- list()
authors <- list()
for (quotes_page in first_quotes_page) {
  print(paste("Processing ", quotes_page))
  new_quotes <- list()
  new_quotes <- lapply(get_quotes_elements(quotes_page), get_quote)
  quotes <- append(quotes, new_quotes)
}

问题原因

在get_quote函数中,你直接对authors执行append操作,但R默认将函数内的变量视为局部变量——即便全局环境存在同名的authors,函数内的authors会被当成新的局部变量创建,不会修改全局的列表。

解决方法

方法1:使用全局变量赋值符

在get_quote函数里,把authors <- append(authors, author)改成:

authors <<- append(authors, author)

<<-赋值符会直接修改全局环境中的authors变量,而非创建局部变量。

方法2:返回作者信息并外部收集(更推荐)

避免依赖全局变量,让get_quote同时返回名言和作者信息,再在外部循环中分别收集:

  1. 修改get_quote函数:
get_quote <- function(quote_element) {
  quote <- list()
  quote_text = html_nodes(quote_element,'.text') %>% html_text()
  quote_author = html_nodes(quote_element,'.author') %>% html_text()
  quote_tags = html_nodes(quote_element,'.tags') %>% html_text()
  quote['Author'] <- quote_author
  quote['Quote'] <- quote_text
  quote['Tags'] <- quote_tags
  author_page_url <- paste0('https://quotes.toscrape.com', html_nodes(quote_element,"a[href*='/author/']") %>% html_attr("href"))
  author <- get_author(author_page_url)
  print(paste(" Quote from ", author["Name"], " added"))
  # 返回包含名言和作者的列表
  return(list(quote = quote, author = author))
}
  1. 修改运行脚本:
quotes_pages <- get_quotes_pages(start_url)
first_quotes_page <- quotes_pages[[1]]
quotes <- list()
authors <- list()
for (quotes_page in first_quotes_page) {
  print(paste("Processing ", quotes_page))
  # 获取包含名言和作者的结果集
  results <- lapply(get_quotes_elements(quotes_page), get_quote)
  # 分别提取并追加到对应列表
  quotes <- append(quotes, lapply(results, function(x) x$quote))
  authors <- append(authors, lapply(results, function(x) x$author))
}

这种方式更符合函数式编程规范,代码模块化更强,也更易调试维护。

额外优化建议

  • get_quotes_pages中硬编码pages_count <- 9不够灵活,建议通过判断页面是否存在「下一页」按钮来循环获取所有页面,避免固定页数导致的遗漏或错误。
  • 已经全局创建了session对象,函数内可以直接用session调用节点,无需每次session_jump_to后再read_html,简化代码。

内容的提问来源于stack exchange,提问作者Didier mac cormick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 22:32:11