如何用R抓取FinViz个股表格关键统计数据?求R爬虫经验指导
从FinViz抓取个股关键统计数据(R语言实现)
Hey there, let's tackle your FinViz scraping question step by step. I’ve got plenty of R web scraping experience under my belt, so let’s fix up your code and make it way more reliable.
1. 先修复你的现有XML代码
First, let's address the issues in your current code:
- You had an unclosed parenthesis in the
assign()call - FinViz blocks requests without a proper User-Agent header
- Using
tableNodes[[9]]relies on fragile page structure indexing (prone to break if FinViz updates their layout)
Here's the fixed version of your original approach:
library(XML) library(httr) stocks <- c("AAPL","MSFT") # Set a valid User-Agent to avoid being blocked headers <- add_headers("User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") for (s in stocks) { url <- paste0("http://finviz.com/quote.ashx?t=", s) # Fetch page with headers webpage <- content(GET(url, headers), as = "text") html <- htmlTreeParse(webpage, useInternalNodes = TRUE, asText = TRUE) # Target the specific stats table using XPath (more reliable than index) tableNodes <- getNodeSet(html, "//table[@class='snapshot-table2']") # Fix the assign syntax and use better headers assign(s, readHTMLTable(tableNodes[[1]], header = FALSE, stringsAsFactors = FALSE)) # Add a delay to avoid rate limiting Sys.sleep(2) } # Check the result for AAPL head(AAPL)
2. 更现代可靠的实现(用rvest包)
The XML package is a bit dated—most R scrapers now use rvest (part of the tidyverse) for cleaner, more maintainable code. Here's a robust implementation:
Step 1: 安装并加载所需包
install.packages(c("rvest", "dplyr", "purrr")) library(rvest) library(dplyr) library(purrr)
Step 2: 定义可复用的抓取函数
scrape_finviz_stats <- function(ticker) { # 设置请求头绕过反爬机制 request_headers <- c( "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" ) # 构建FinViz个股页面URL finviz_url <- paste0("https://finviz.com/quote.ashx?t=", ticker) # 获取并解析页面 page_html <- read_html(finviz_url, httr::add_headers(.headers = request_headers)) # 用CSS选择器定位核心统计表格(比索引更稳定) stats_table <- page_html %>% html_element(css = ".snapshot-table2") %>% html_table(header = FALSE) # 将表格整理为清晰的键值对格式 clean_stats <- stats_table %>% select(Metric = X1, Value = X2) %>% mutate(Ticker = ticker) return(clean_stats) }
Step 3: 批量抓取多个个股数据
# 要抓取的个股代码列表 target_tickers <- c("AAPL", "MSFT") # 抓取所有个股并合并为单个数据框 all_stock_stats <- map_dfr(target_tickers, scrape_finviz_stats) # 预览结果 head(all_stock_stats)
3. 关键注意事项
- 避免被封禁: 在请求之间添加
Sys.sleep(2)来遵守FinViz的速率限制,不要连续抓取上百只个股而不设延迟。 - 适配页面更新: 如果FinViz改版,用浏览器开发者工具(F12)找到统计表格的新CSS选择器。
- 清洗数值数据: 抓取到的数值常带有
%、B或M等后缀,这里提供一个快速转换为数字的方法:
clean_numeric_stats <- all_stock_stats %>% mutate( Numeric_Value = case_when( grepl("%$", Value) ~ as.numeric(sub("%", "", Value)) / 100, grepl("B$", Value) ~ as.numeric(sub("B", "", Value)) * 1e9, grepl("M$", Value) ~ as.numeric(sub("M", "", Value)) * 1e6, TRUE ~ as.numeric(Value) ) )
内容的提问来源于stack exchange,提问作者Frank
相关产品推荐
相关产品推荐

