You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中从API隐藏网页提取JSON:Liberty Shares数据获取求助

Troubleshooting Liberty Shares Data Scraping Issues

Let me walk through what might be tripping up each of your methods and share actionable fixes to get that target table data:

1. Why rvest Only Loads the Page Header

Static HTML parsers like read_html() can’t handle content loaded dynamically via JavaScript (e.g., AJAX calls, React/Vue rendering). The table you’re after is almost certainly being injected into the page after the initial HTML loads, which is why your static parse only grabs the header.

Fix: Use a browser automation tool to fully render the page first, then parse the loaded content with rvest. Here’s an example with RSelenium:

library(RSelenium)
library(rvest)

# Start a headless Chrome session
driver <- rsDriver(browser = "chrome", chromever = "latest", extraCapabilities = list(chromeOptions = list(args = c("--headless=new"))))
remDr <- driver[["client"]]

# Navigate to your target Liberty Shares page
remDr$navigate("https://libertyshares.example.com/your-target-page") # Replace with actual URL

# Wait for the table to load (adjust the CSS selector to match your table's ID/class)
remDr$waitForElement(using = "css", value = "#target-table", timeout = 10000)

# Grab the fully rendered page source
page_source <- remDr$getPageSource()[[1]]
html <- read_html(page_source)

# Extract the table like you normally would with rvest
target_table <- html %>% html_element("#target-table") %>% html_table()

# Clean up the browser session
remDr$close()
driver$server$stop()

2. PhantomJS Failing to Capture the Table

PhantomJS is outdated and no longer maintained—many modern websites block or fail to render correctly for it. Even if you get it running, it might not wait long enough for the table’s dynamic content to load.

Fix: Switch to a modern headless browser like Chrome or Firefox. For R, the playwright package is a great, well-supported option:

library(playwright)

# Initialize playwright with headless Chrome
pw <- playwright$launch()
browser <- pw$chromium$launch(headless = TRUE)
page <- browser$newPage()

# Navigate and wait for the table to load
page$goto("https://libertyshares.example.com/your-target-page")
page$waitForSelector("#target-table") # Use your table's specific selector

# Extract the fully rendered page content and parse
html <- read_html(page$content())
target_table <- html %>% html_element("#target-table") %>% html_table()

# Clean up
browser$close()
pw$stop()

3. Issues with XHR/Curl Requests

If the table data loads via an XHR call, missing critical request details (like headers, cookies, or dynamic parameters) will cause your curl/R request to fail.

Steps to Fix:

  1. Re-open your browser’s Network tab, filter for XHR/fetch requests, and locate the one that loads the table data.
  2. Right-click the request and copy all relevant details:
    • Request method (GET/POST)
    • Headers (pay close attention to User-Agent, Referer, Authorization, or any custom tokens)
    • Query parameters or POST body
    • Cookies (if the site requires a session or authentication)
  3. Use the httr package in R to replicate the request exactly. Example for a GET request with custom headers:
library(httr)

response <- GET(
  url = "https://libertyshares.example.com/api/table-data", # Replace with actual API endpoint
  add_headers(
    `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    `Referer` = "https://libertyshares.example.com/your-target-page",
    `Authorization` = "Bearer YOUR_TOKEN_IF_NEEDED"
  ),
  query = list(
    param1 = "value1",
    param2 = "value2" # Add any query params from the XHR request
  ),
  set_cookies(
    cookie_name = "cookie_value" # Add required cookies here
  )
)

# Parse the JSON response into a usable format
table_data <- content(response, "parsed")
  1. If you copied a curl command, use the curlconverter package to convert it directly to R code:
library(curlconverter)

# Paste your copied curl command here
curl_cmd <- 'curl "https://libertyshares.example.com/api/table-data" -H "User-Agent: ..."'
req <- make_req(straighten(curl_cmd)[[1]])
response <- req()
table_data <- content(response, "parsed")

Final Notes

Always check the website’s robots.txt and terms of service to ensure scraping is allowed. If you hit anti-scraping measures (like rate limits or bot detection), add delays between requests or rotate user agents to avoid being blocked.

内容的提问来源于stack exchange,提问作者GonzaloXavier

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:56:30