使用R语言rvest包爬取DailyFX情绪数据返回空值如何解决?
问题原因
你使用的选择器逻辑正确,但目标站点的持仓数据为JavaScript动态渲染生成,rvest的read_html()仅能抓取页面初始静态HTML内容,不会执行JS代码加载后续渲染的元素,因此匹配不到目标节点,返回character(0)。
解决方案
方案1:使用RSelenium模拟浏览器加载(兼容原选择器逻辑)
该方案通过模拟真实浏览器访问,等待JS渲染完成后再提取数据,可直接复用你之前通过Selector Gadget获取的选择器:
- 前置准备:安装
RSelenium包,且提前下载对应你Chrome浏览器版本的ChromeDriver,配置到系统环境变量 - 示例代码:
library(RSelenium) library(rvest) # 启动Chrome驱动,chromever参数替换为你本地Chrome的版本号 rd <- rsDriver(browser = "chrome", chromever = "xxx.xx.xxxx.xx", port = 4567L) remDr <- rd$client # 访问目标页面,等待渲染完成 remDr$navigate("https://www.dailyfx.com/sentiment") Sys.sleep(3) # 获取渲染后的完整页面源码 page_source <- remDr$getPageSource()[[1]] wp <- read_html(page_source) # 复用原选择器提取数据 bal <- html_nodes(wp, ".dfx-technicalSentimentCard--opened .dfx-technicalSentimentCard__netLongContainer .font-weight-bold") bal <- html_text(bal) print(bal) # 用完关闭驱动释放资源 remDr$close() rd$server$stop()
方案2:提取页面内嵌JSON数据(轻量高效,无需浏览器)
目标站点的初始持仓数据直接内嵌在静态页面的JSON结构中,不需要模拟浏览器即可直接提取,运行效率更高:
- 前置准备:安装
jsonlite包用于解析JSON - 示例代码:
library(rvest) library(jsonlite) link <- "https://www.dailyfx.com/sentiment" wp <- read_html(link) # 提取页面内嵌的全量数据 raw_data <- html_elements(wp, "#__NEXT_DATA__") %>% html_text() parsed_data <- fromJSON(raw_data) full_sentiment <- parsed_data$props$pageProps$technicalSentimentState$sentimentData # 提取默认展开的EUR/USD净多头持仓比例,如需其他产品可替换对应的产品名称 eurusd_net_long <- full_sentiment$`EUR/USD`$net_long_percent print(eurusd_net_long)
内容的提问来源于stack exchange,提问作者mr.T
相关产品推荐
相关产品推荐

