Rcrawler v0.1.7爬取部分HTTP/HTTPS网站失败问题求助
Hey there, let’s work through this Rcrawler problem together. It’s frustrating when most sites work but a handful don’t—even switching between HTTP and HTTPS doesn’t fix it. Here are practical steps to diagnose and resolve the issue:
Check for anti-scraping measures
A lot of sites block default crawler user agents. Try overriding Rcrawler’s default UA with a browser-like string to avoid being flagged:Rcrawler(Url = "your_problem_url", UserAgent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")Verify network and proxy settings
Sometimes local firewalls, corporate proxies, or network restrictions block crawler requests. First, confirm you can access the problematic sites in a regular browser. If you need a proxy to reach them, configure Rcrawler to use it:Rcrawler(Url = "your_problem_url", Proxy = "http://your_proxy_address:port")Check robots.txt rules
Some sites explicitly prohibit crawlers via theirrobots.txtfile (e.g.,https://example.com/robots.txt). If the site blocks crawlers, you can bypass this check (use responsibly, only if you have legal access):Rcrawler(Url = "your_problem_url", RobotsTxt = FALSE)Dig into error logs
Enable logging to get specific details about why requests fail. This will show you if it’s a timeout, 403 Forbidden, or another issue:crawl_result <- Rcrawler(Url = "your_problem_url", log = TRUE) # View the error log for the failed request print(crawl_result$ErrorLog)Test single sites first
Batch crawling can trigger anti-scraping systems faster. Isolate one problematic URL and test it alone to rule out rate-limiting or bulk-request blocks.Update Rcrawler to the latest version
Even v0.1.7 claims HTTPS support, recent updates might fix edge-case bugs with both HTTP and HTTPS sites. Install the latest dev version with:devtools::install_github("salimk/Rcrawler")
内容的提问来源于stack exchange,提问作者amarbut

