如何使用R语言爬取Finviz网站的季度财务报表表格
问题原因说明
你用rvest无法抓取三类财报表格的核心原因是:finviz的季度利润表、资产负债表、现金流量表属于客户端动态加载内容,页面初始静态HTML返回时不包含这部分数据,是浏览器加载完页面后异步调用后端接口拉取数据再渲染到页面的,rvest仅支持解析静态HTML,自然抓不到这部分内容。
可行解决方案(R语言环境)
以下两个方案均可落地,可根据自身需求选择:
方案1:直接调用财报数据接口(推荐,效率最高,无需模拟浏览器)
finviz的个股财报接口无访问权限限制,格式固定,返回结构化HTML片段可直接用rvest解析:
- 接口规则:
https://finviz.com/api/statement.ashx?t=股票代码&statement=类型&period=quarterly
其中类型参数对应三类报表:IA:利润表BA:资产负债表CA:现金流量表
- 示例代码:
library(rvest) library(httr) # 配置请求头模拟浏览器,避免被反爬拦截 headers <- c( "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Referer" = "https://finviz.com/" ) # 单只股票爬取示例,可自行扩展为多股票循环 ticker <- "A" statements <- list(income = "IA", balance = "BA", cashflow = "CA") result <- list() for (name in names(statements)) { req_url <- paste0("https://finviz.com/api/statement.ashx?t=", ticker, "&statement=", statements[[name]], "&period=quarterly") resp <- GET(req_url, add_headers(.headers = headers)) result[[name]] <- read_html(content(resp, "text")) %>% html_element("table") %>% html_table(header = TRUE) } # 输出对应报表查看 print(result$income) print(result$balance) print(result$cashflow)
- 批量爬取注意事项:
- 每次请求间隔设置1-2秒,避免触发频率限制
- 出现403报错时可更换User-Agent或添加代理池解决
方案2:使用RSelenium模拟浏览器加载(适合少量数据爬取场景)
如果不想分析接口规则,可通过RSelenium启动真实浏览器加载完整页面后提取表格:
- 示例代码片段:
library(RSelenium) library(rvest) # 启动Chrome,需提前安装对应版本的ChromeDriver driver <- rsDriver(browser = "chrome", chromever = "你本地Chrome的版本号")$client driver$navigate("https://finviz.com/quote.ashx?t=A&ty=c&p=d&b=1") # 等待动态内容加载完成 Sys.sleep(3) page_html <- read_html(driver$getPageSource()[[1]]) # 按标题定位对应报表 income_table <- page_html %>% html_element(xpath = "//h2[text()='Income Statement']/following-sibling::table[1]") %>% html_table(header = TRUE) # 资产负债表、现金流量表修改xpath中的标题文本即可提取
- 注意:该方案运行效率远低于接口方案,不适合大规模批量爬取使用。
内容的提问来源于stack exchange,提问作者Citizen
相关产品推荐
相关产品推荐

