You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R语言爬取《华尔街日报》2024年400所大学排名数据?

爬取WSJ 2024最佳大学排名的R语言实操方案

核心问题分析

你之前用html_elements("section") %>% html_element("p")定位太宽泛,WSJ的排名数据封装在带特定标识的卡片容器里,得用更精准的选择器定位。下面是分步解决方法:


1. 加载依赖包

除了rvest,搭配dplyr和stringr处理数据更高效:

library(rvest)
library(dplyr)
library(stringr)

2. 定位排名卡片容器

WSJ给排名条目标注了data-test-id专属标识,直接抓取所有排名卡片:

link <- "https://www.wsj.com/rankings/college-rankings/best-colleges-2024"
wsj <- read_html(link)

# 抓取所有排名卡片(可通过浏览器右键「检查元素」确认选择器有效性)
rank_cards <- wsj %>% html_elements("div[data-test-id='college-ranking-card']")

3. 提取单条排名数据

从每个卡片中提取排名、大学名称、得分等关键信息:

# 提取排名
ranks <- rank_cards %>% 
  html_element("span[data-test-id='rank']") %>% 
  html_text() %>% 
  as.integer()

# 提取大学名称
college_names <- rank_cards %>% 
  html_element("h3[data-test-id='college-name']") %>% 
  html_text()

# 提取总体得分(清洗文本,仅保留数值部分)
overall_scores <- rank_cards %>% 
  html_element("div[data-test-id='overall-score']") %>% 
  html_text() %>% 
  str_extract("\\d+\\.\\d+") %>% 
  as.numeric()

# 提取毕业后薪资指标(示例,可根据需求添加教学质量、录取率等其他指标)
salary_metrics <- rank_cards %>% 
  html_element("div[data-test-id='outcome-salary']") %>% 
  html_text()

4. 组合成数据集

把提取的字段整合成结构化数据框:

wsj_college_rankings <- tibble(
  排名 = ranks,
  大学名称 = college_names,
  总体得分 = overall_scores,
  毕业后薪资 = salary_metrics
)

# 查看前10条数据验证结果
head(wsj_college_rankings, 10)

关键注意事项

  • 动态加载问题:当前页面默认仅显示部分排名(如前50条),要获取全部400条,需用RSelenium或playwright模拟浏览器滚动加载所有内容后再抓取。
  • 反爬限制:WSJ有反爬机制,不要频繁请求,建议每操作一次添加Sys.sleep(2)的延迟。
  • 选择器更新:网站可能调整HTML结构,若代码失效,右键页面元素→「检查」,确认最新的data-test-id或类名。

内容的提问来源于stack exchange,提问作者androsrj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 04:07:09