如何用R编程获取网页内容大小及维基百科高尔夫球手爬取优化
关于R语言获取网页内容大小及维基百科页面有效性判断的解决方案
让我一步步帮你解决这两个问题——先搞定R语言获取网页内容大小的方法,再聊聊爬维基百科高尔夫球手页面的更优方案~
一、如何用R编程获取网页内容大小?
获取网页内容大小有两种实用思路,选哪种取决于你是否需要下载完整页面:
1. 用HEAD请求读取响应头(高效,无需下载内容)
HEAD请求只返回HTTP响应头,不会传输页面主体,速度更快,适合仅需获取大小的场景。用httr包就能轻松实现:
library(httr) get_content_size_head <- function(url) { # 自定义User-Agent,避免被网站拦截 response <- HEAD(url, user_agent("My Web Scraper (your.email@example.com)")) # 从响应头提取Content-Length并转换为KB content_length <- as.numeric(headers(response)[["Content-Length"]]) # 处理动态页面(无Content-Length的情况) if (is.na(content_length)) { warning("该页面未返回Content-Length,可能是动态生成页面") return(NA) } return(content_length / 1024) } # 示例调用 get_content_size_head("https://en.wikipedia.org/wiki/Tiger_Woods")
2. 下载页面后计算大小(更准确,适配动态页面)
很多JS渲染的动态页面不会返回Content-Length,这时需要用GET请求下载内容,再计算字节大小:
get_content_size_get <- function(url) { response <- GET(url, user_agent("My Web Scraper (your.email@example.com)")) # 确保请求成功,失败则抛出错误 stop_for_status(response) # 获取原始内容的字节数并转换为KB content_size <- length(content(response, as = "raw")) return(content_size / 1024) } # 示例调用 get_content_size_get("https://en.wikipedia.org/wiki/Tiger_Woods")
二、爬取维基百科高尔夫球手页面:更优的页面有效性判断方案
你原本想用内容大小判断页面是否有效,这个思路有明显漏洞:维基百科的404页面或消歧义页可能也会超过2KB,而某些简短的有效球员页面可能更小,很容易误判。推荐两种更可靠的方案:
1. 检查HTTP状态码(最直接的判断方式)
维基百科对不存在的页面会返回404 Not Found,有效页面则返回200 OK。我们可以利用这一点做判断:
library(httr) library(rvest) get_golfer_page <- function(golfer_name) { # 生成两个候选URL,注意用URLencode处理特殊字符 url_with_suffix <- paste0("https://en.wikipedia.org/wiki/", URLencode(paste0(golfer_name, "(golfer)"))) url_original <- paste0("https://en.wikipedia.org/wiki/", URLencode(golfer_name)) # 优先尝试带(golfer)后缀的URL response <- GET(url_with_suffix, user_agent("My Golf Scraper (your.email@example.com)")) if (http_status(response)$category == "Success") { message("成功获取带后缀的页面:", golfer_name) return(read_html(response)) } # 后缀URL无效,尝试原始姓名URL response <- GET(url_original, user_agent("My Golf Scraper (your.email@example.com)")) if (http_status(response)$category == "Success") { message("成功获取原始页面:", golfer_name) return(read_html(response)) } # 两个URL都无效 warning("无法获取", golfer_name, "的页面") return(NULL) } # 示例调用:测试重名球员(比如"John Smith") page <- get_golfer_page("John Smith")
2. 结合页面内容特征(应对消歧义页)
有时候原始URL对应的是消歧义页(状态码200,但不是目标球员页面),这时可以检查页面的标题特征:
# 判断是否为消歧义页的辅助函数 is_disambiguation_page <- function(html_page) { # 维基百科消歧义页的标题通常包含"Disambiguation" page_title <- html_text(html_element(html_page, "h1")) return(grepl("Disambiguation", page_title, ignore.case = TRUE)) } # 改进后的页面获取函数 get_golfer_page_improved <- function(golfer_name) { url_with_suffix <- paste0("https://en.wikipedia.org/wiki/", URLencode(paste0(golfer_name, "(golfer)"))) url_original <- paste0("https://en.wikipedia.org/wiki/", URLencode(golfer_name)) # 尝试后缀URL response <- GET(url_with_suffix, user_agent("My Golf Scraper (your.email@example.com)")) if (http_status(response)$category == "Success") { html <- read_html(response) if (!is_disambiguation_page(html)) { message("成功获取带后缀的球员页面:", golfer_name) return(html) } } # 尝试原始URL response <- GET(url_original, user_agent("My Golf Scraper (your.email@example.com)")) if (http_status(response)$category == "Success") { html <- read_html(response) if (!is_disambiguation_page(html)) { message("成功获取原始球员页面:", golfer_name) return(html) } else { warning("原始页面是消歧义页,请确认球员姓名:", golfer_name) } } warning("无法获取有效的球员页面:", golfer_name) return(NULL) }
重要提醒
爬取维基百科时,一定要设置自定义的User-Agent(不要用默认的httr标识),并遵守维基百科的爬虫规范,避免短时间内频繁发送请求导致IP被封禁。
内容的提问来源于stack exchange,提问作者pssguy
相关产品推荐
相关产品推荐

