You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R编程获取网页内容大小及维基百科高尔夫球手爬取优化

关于R语言获取网页内容大小及维基百科页面有效性判断的解决方案

让我一步步帮你解决这两个问题——先搞定R语言获取网页内容大小的方法,再聊聊爬维基百科高尔夫球手页面的更优方案~

一、如何用R编程获取网页内容大小?

获取网页内容大小有两种实用思路,选哪种取决于你是否需要下载完整页面:

1. 用HEAD请求读取响应头(高效,无需下载内容)

HEAD请求只返回HTTP响应头,不会传输页面主体,速度更快,适合仅需获取大小的场景。用httr包就能轻松实现:

library(httr)

get_content_size_head <- function(url) {
  # 自定义User-Agent,避免被网站拦截
  response <- HEAD(url, user_agent("My Web Scraper (your.email@example.com)"))
  # 从响应头提取Content-Length并转换为KB
  content_length <- as.numeric(headers(response)[["Content-Length"]])
  # 处理动态页面(无Content-Length的情况)
  if (is.na(content_length)) {
    warning("该页面未返回Content-Length,可能是动态生成页面")
    return(NA)
  }
  return(content_length / 1024)
}

# 示例调用
get_content_size_head("https://en.wikipedia.org/wiki/Tiger_Woods")

2. 下载页面后计算大小(更准确,适配动态页面)

很多JS渲染的动态页面不会返回Content-Length,这时需要用GET请求下载内容,再计算字节大小:

get_content_size_get <- function(url) {
  response <- GET(url, user_agent("My Web Scraper (your.email@example.com)"))
  # 确保请求成功,失败则抛出错误
  stop_for_status(response)
  # 获取原始内容的字节数并转换为KB
  content_size <- length(content(response, as = "raw"))
  return(content_size / 1024)
}

# 示例调用
get_content_size_get("https://en.wikipedia.org/wiki/Tiger_Woods")

二、爬取维基百科高尔夫球手页面:更优的页面有效性判断方案

你原本想用内容大小判断页面是否有效,这个思路有明显漏洞:维基百科的404页面或消歧义页可能也会超过2KB,而某些简短的有效球员页面可能更小,很容易误判。推荐两种更可靠的方案:

1. 检查HTTP状态码(最直接的判断方式)

维基百科对不存在的页面会返回404 Not Found,有效页面则返回200 OK。我们可以利用这一点做判断:

library(httr)
library(rvest)

get_golfer_page <- function(golfer_name) {
  # 生成两个候选URL,注意用URLencode处理特殊字符
  url_with_suffix <- paste0("https://en.wikipedia.org/wiki/", URLencode(paste0(golfer_name, "(golfer)")))
  url_original <- paste0("https://en.wikipedia.org/wiki/", URLencode(golfer_name))
  
  # 优先尝试带(golfer)后缀的URL
  response <- GET(url_with_suffix, user_agent("My Golf Scraper (your.email@example.com)"))
  if (http_status(response)$category == "Success") {
    message("成功获取带后缀的页面:", golfer_name)
    return(read_html(response))
  }
  
  # 后缀URL无效,尝试原始姓名URL
  response <- GET(url_original, user_agent("My Golf Scraper (your.email@example.com)"))
  if (http_status(response)$category == "Success") {
    message("成功获取原始页面:", golfer_name)
    return(read_html(response))
  }
  
  # 两个URL都无效
  warning("无法获取", golfer_name, "的页面")
  return(NULL)
}

# 示例调用:测试重名球员(比如"John Smith")
page <- get_golfer_page("John Smith")

2. 结合页面内容特征(应对消歧义页)

有时候原始URL对应的是消歧义页(状态码200,但不是目标球员页面),这时可以检查页面的标题特征:

# 判断是否为消歧义页的辅助函数
is_disambiguation_page <- function(html_page) {
  # 维基百科消歧义页的标题通常包含"Disambiguation"
  page_title <- html_text(html_element(html_page, "h1"))
  return(grepl("Disambiguation", page_title, ignore.case = TRUE))
}

# 改进后的页面获取函数
get_golfer_page_improved <- function(golfer_name) {
  url_with_suffix <- paste0("https://en.wikipedia.org/wiki/", URLencode(paste0(golfer_name, "(golfer)")))
  url_original <- paste0("https://en.wikipedia.org/wiki/", URLencode(golfer_name))
  
  # 尝试后缀URL
  response <- GET(url_with_suffix, user_agent("My Golf Scraper (your.email@example.com)"))
  if (http_status(response)$category == "Success") {
    html <- read_html(response)
    if (!is_disambiguation_page(html)) {
      message("成功获取带后缀的球员页面:", golfer_name)
      return(html)
    }
  }
  
  # 尝试原始URL
  response <- GET(url_original, user_agent("My Golf Scraper (your.email@example.com)"))
  if (http_status(response)$category == "Success") {
    html <- read_html(response)
    if (!is_disambiguation_page(html)) {
      message("成功获取原始球员页面:", golfer_name)
      return(html)
    } else {
      warning("原始页面是消歧义页,请确认球员姓名:", golfer_name)
    }
  }
  
  warning("无法获取有效的球员页面:", golfer_name)
  return(NULL)
}

重要提醒

爬取维基百科时,一定要设置自定义的User-Agent(不要用默认的httr标识),并遵守维基百科的爬虫规范,避免短时间内频繁发送请求导致IP被封禁。

内容的提问来源于stack exchange,提问作者pssguy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:58:27