You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用purrr::map替代for循环批量抓取小说站文章标题?

用purrr::map替代for循环优化R语言网页抓取代码

优化后的map实现代码

library(rvest)
library(purrr)

# 读取目标网页HTML
html_narou <- read_html("https://yomou.syosetu.com/search.php?&type=er&order_former=search&order=new&notnizi=1&p=1",
                       encoding = "UTF-8")

# 生成所有需要的CSS选择器
base_css_former <- "#main_search > div:nth-child("
base_css_latter <- ") > div > a"
css_selectors <- paste0(base_css_former, 5:24, base_css_latter)

# 用map批量提取标题文本
narou_titles <- map(css_selectors, ~{
  html_element(html_narou, css = .x) %>% 
    html_text()
}) %>% unlist()

代码说明

  • 先一次性生成所有目标CSS选择器的列表,避免循环内重复拼接字符串
  • purrr::map会遍历css_selectors中的每一个选择器,执行括号内的提取操作,返回一个包含所有标题的列表
  • 最后用unlist()把列表转换为向量,和原for循环输出的格式完全一致

更高效的简化方案

其实不需要循环或map,直接用更精准的CSS选择器就能一步提取所有目标标题,效率更高:

library(rvest)

html_narou <- read_html("https://yomou.syosetu.com/search.php?&type=er&order_former=search&order=new&notnizi=1&p=1",
                       encoding = "UTF-8")

# 用:gt(3)匹配#main_search下索引大于3的div(对应原代码的5-24位元素)
narou_titles <- html_elements(html_narou, css = "#main_search > div:gt(3) > div > a") %>% 
  html_text()

内容的提问来源于stack exchange,提问作者Ashu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 06:50:28