You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R爬取Skysports网站EFL Championship赛程表的参数疑问

用R爬取Skysports英冠赛程表的解决方案

所需依赖包

首先确保安装并加载以下工具包:

# 首次运行时安装包
install.packages(c("rvest", "dplyr", "writexl"))

# 加载包
library(rvest)
library(dplyr)
library(writexl)

完整爬取代码

# 目标赛程页面URL
url <- "https://www.skysports.com/championship-fixtures"

# 读取页面HTML内容
page <- read_html(url)

# 定位所有单场赛程的父节点
fixture_nodes <- page %>% html_nodes(".fixres__item")

# 批量提取每场赛程的关键信息
fixtures_df <- lapply(fixture_nodes, function(node) {
  # 提取比赛日期
  match_date <- node %>% html_node(".fixres__header2") %>% html_text(trim = TRUE)
  # 提取主队名称
  home_team <- node %>% html_node(".fixres__team--home .swap-text__target") %>% html_text(trim = TRUE)
  # 提取客队名称
  away_team <- node %>% html_node(".fixres__team--away .swap-text__target") %>% html_text(trim = TRUE)
  # 提取开球时间
  kickoff_time <- node %>% html_node(".fixres__status") %>% html_text(trim = TRUE)
  
  # 整合成数据框行
  data.frame(
    Date = match_date,
    Home_Team = home_team,
    Away_Team = away_team,
    Kickoff_Time = kickoff_time,
    stringsAsFactors = FALSE
  )
}) %>% bind_rows()

# 导出为Excel文件
write_xlsx(fixtures_df, "EFL_Championship_Fixtures.xlsx")

参数说明(如何确定html_nodes/html_node的选择器)

这些选择器都是通过浏览器开发者工具(按F12打开)分析页面HTML结构得到的:

  • .fixres__item:包裹单场完整赛程的容器节点,是所有赛程项的共同父元素
  • .fixres__header2:单场赛程对应的日期文本节点
  • .fixres__team--home .swap-text__target:主队名称的具体文本节点(swap-text__target是Skysports用来显示队名的元素类)
  • .fixres__team--away .swap-text__target:客队名称的具体文本节点
  • .fixres__status:开球时间的文本节点

注意事项

  • 网站的HTML结构可能随时间更新,若后续爬取失败,需重新用开发者工具检查元素类名是否变化
  • 爬取频率不要过高,避免触发网站反爬机制

内容的提问来源于stack exchange,提问作者user12025050

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 04:32:43