在R中从API隐藏网页提取JSON:Liberty Shares数据获取求助
Let me walk through what might be tripping up each of your methods and share actionable fixes to get that target table data:
1. Why rvest Only Loads the Page Header
Static HTML parsers like read_html() can’t handle content loaded dynamically via JavaScript (e.g., AJAX calls, React/Vue rendering). The table you’re after is almost certainly being injected into the page after the initial HTML loads, which is why your static parse only grabs the header.
Fix: Use a browser automation tool to fully render the page first, then parse the loaded content with rvest. Here’s an example with RSelenium:
library(RSelenium) library(rvest) # Start a headless Chrome session driver <- rsDriver(browser = "chrome", chromever = "latest", extraCapabilities = list(chromeOptions = list(args = c("--headless=new")))) remDr <- driver[["client"]] # Navigate to your target Liberty Shares page remDr$navigate("https://libertyshares.example.com/your-target-page") # Replace with actual URL # Wait for the table to load (adjust the CSS selector to match your table's ID/class) remDr$waitForElement(using = "css", value = "#target-table", timeout = 10000) # Grab the fully rendered page source page_source <- remDr$getPageSource()[[1]] html <- read_html(page_source) # Extract the table like you normally would with rvest target_table <- html %>% html_element("#target-table") %>% html_table() # Clean up the browser session remDr$close() driver$server$stop()
2. PhantomJS Failing to Capture the Table
PhantomJS is outdated and no longer maintained—many modern websites block or fail to render correctly for it. Even if you get it running, it might not wait long enough for the table’s dynamic content to load.
Fix: Switch to a modern headless browser like Chrome or Firefox. For R, the playwright package is a great, well-supported option:
library(playwright) # Initialize playwright with headless Chrome pw <- playwright$launch() browser <- pw$chromium$launch(headless = TRUE) page <- browser$newPage() # Navigate and wait for the table to load page$goto("https://libertyshares.example.com/your-target-page") page$waitForSelector("#target-table") # Use your table's specific selector # Extract the fully rendered page content and parse html <- read_html(page$content()) target_table <- html %>% html_element("#target-table") %>% html_table() # Clean up browser$close() pw$stop()
3. Issues with XHR/Curl Requests
If the table data loads via an XHR call, missing critical request details (like headers, cookies, or dynamic parameters) will cause your curl/R request to fail.
Steps to Fix:
- Re-open your browser’s Network tab, filter for XHR/fetch requests, and locate the one that loads the table data.
- Right-click the request and copy all relevant details:
- Request method (GET/POST)
- Headers (pay close attention to
User-Agent,Referer,Authorization, or any custom tokens) - Query parameters or POST body
- Cookies (if the site requires a session or authentication)
- Use the
httrpackage in R to replicate the request exactly. Example for a GET request with custom headers:
library(httr) response <- GET( url = "https://libertyshares.example.com/api/table-data", # Replace with actual API endpoint add_headers( `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", `Referer` = "https://libertyshares.example.com/your-target-page", `Authorization` = "Bearer YOUR_TOKEN_IF_NEEDED" ), query = list( param1 = "value1", param2 = "value2" # Add any query params from the XHR request ), set_cookies( cookie_name = "cookie_value" # Add required cookies here ) ) # Parse the JSON response into a usable format table_data <- content(response, "parsed")
- If you copied a curl command, use the
curlconverterpackage to convert it directly to R code:
library(curlconverter) # Paste your copied curl command here curl_cmd <- 'curl "https://libertyshares.example.com/api/table-data" -H "User-Agent: ..."' req <- make_req(straighten(curl_cmd)[[1]]) response <- req() table_data <- content(response, "parsed")
Final Notes
Always check the website’s robots.txt and terms of service to ensure scraping is allowed. If you hit anti-scraping measures (like rate limits or bot detection), add delays between requests or rotate user agents to avoid being blocked.
内容的提问来源于stack exchange,提问作者GonzaloXavier

