使用R语言读取纳斯达克FB页面时出现全局冻结问题求助
Hey there, let’s dig into this frustrating freeze issue you’re facing. For months you could pull data from the NASDAQ FB page no problem, but starting Wednesday every method you try locks up—across both Windows R Studio and Ubuntu R environments. Let’s start by restating the problem clearly, then walk through actionable fixes.
The Problem Recap
You’ve been using code like this successfully until recently:
myURL <- "http://www.nasdaq.com/symbol/fb" webpage <- readLines(myURL)
Now, every approach you’ve tried freezes:
- Base R’s
readLines()(previously reliable) rvestfunctions:read_html(),html_session()(even after resetting the user agent)httr’sGET()RCurl’sgetURL()
You’ve already done initial checks in Chrome, so let’s focus on R-specific fixes and common anti-scraping roadblocks.
Likely Cause & Fixes
The most probable culprit here is that NASDAQ updated their anti-scraping defenses—this is super common when sites notice automated traffic. Here’s what you can try:
1. Mimic a Real Browser Fully
Just changing the user agent isn’t enough anymore. Modern sites check a whole set of request headers to spot bots. Let’s replicate the headers Chrome sends with httr:
library(httr) # These headers are pulled from a real Chrome request browser_headers <- add_headers( `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", `Accept` = "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", `Accept-Language` = "en-US,en;q=0.5", `Referer` = "https://www.nasdaq.com/", `DNT` = "1", `Connection` = "keep-alive", `Upgrade-Insecure-Requests` = "1" ) # Try fetching the HTTPS version (most sites redirect HTTP to HTTPS now) response <- GET("https://www.nasdaq.com/symbol/fb", browser_headers) # Check if we get a valid response http_status(response) content(response, "text")
Note: I switched to the HTTPS URL—many older HTTP requests get stuck on redirects, which can cause freezes.
2. Handle JavaScript-Rendered Content
If NASDAQ now loads page data with JavaScript, basic HTTP requests (like readLines or read_html) can’t access it because they don’t execute JS. Try using a headless browser tool like RSelenium to simulate a real browser:
# Example with RSelenium (you'll need ChromeDriver installed first) library(RSelenium) # Start a Chrome session driver <- rsDriver(browser = "chrome", port = 4567L) remDr <- driver[["client"]] # Navigate to the page and wait for JS to load remDr$navigate("https://www.nasdaq.com/symbol/fb") Sys.sleep(3) # Give time for content to load # Grab the page source page_source <- remDr$getPageSource()[[1]] # Clean up properly remDr$close() driver$server$stop()
3. Rule Out IP Blocking or Rate Limiting
If you’ve been scraping frequently, NASDAQ might have blocked your IP. Try these quick checks:
- Add delays between requests with
Sys.sleep(5)(give the server a break) - Test from a different IP (like a phone hotspot) to see if the freeze goes away
- Use
httr::http_status(response)to check for 403 (forbidden) or 429 (too many requests) status codes
4. Test Outside R First
To confirm the issue isn’t with your network or the site itself, run this curl command in your terminal (Windows Command Prompt or Ubuntu shell):
curl -v https://www.nasdaq.com/symbol/fb
Look for error messages, timeouts, or redirects. If curl also freezes, the problem is likely with the site or your network, not R.
A Quick Reminder
Before you keep scraping, double-check NASDAQ’s terms of service—many sites prohibit automated data collection. If you need regular access, look into their official API (if available) to avoid these issues long-term.
内容的提问来源于stack exchange,提问作者Preston C. Smith

