R语言爬取Google评论时遇Chrome无法访问错误求助
Hey there, let's tackle that frustrating "Chrome not reachable" error you're hitting while trying to scrape Google reviews with RSelenium. I've dealt with this exact issue a handful of times, so here are the most reliable fixes that should get your script back on track:
Common Causes & Solutions
Mismatched Chrome and ChromeDriver Versions
This is the #1 culprit. RSelenium'srsDriverrequires your Chrome browser version to match the ChromeDriver version it uses. First, check your Chrome version (go to Chrome Settings > About Chrome). Then, specify the matching ChromeDriver version when initializingrsDriver:# Replace with your Chrome version number rmDr <- rsDriver(browser = "chrome", chromever = "114.0.5735.90")If you're unsure which version to use, you can use
wdman::chrome_driver()to automatically fetch the correct driver for your installed Chrome.Residual Chrome Processes Blocking the Port
Sometimes leftover Chrome processes (even hidden ones) hog the port RSelenium tries to use. Kill all Chrome instances first:- On Windows: Run
taskkill /im chrome.exe /fin Command Prompt (or addsystem("taskkill /im chrome.exe /f", ignore.stdout = TRUE)to your R script) - On Mac/Linux: Run
pkill chromein Terminal
Then restart your R script.
- On Windows: Run
Port Conflict
The default port RSelenium uses might be occupied by another app. Try specifying a custom, unused port:rmDr <- rsDriver(browser = "chrome", port = 4567L) # Use any number between 1024-65535Outdated Packages
Make sure RSelenium, rvest, and xml2 are all up to date. Reinstall them to be safe:devtools::install_github("ropensci/RSelenium") install.packages(c("rvest", "xml2"))Try Headless Mode
Sometimes graphical Chrome has compatibility issues. Switch to headless mode (no browser window pops up) which is often more stable:# Set headless Chrome options eCaps <- list(chromeOptions = list(args = c('--headless=new', '--disable-gpu'))) rmDr <- rsDriver(browser = "chrome", extraCapabilities = eCaps)
Test Script Example
Here's a cleaned-up version of your code incorporating these fixes:
library(RSelenium) library(rvest) library(xml2) # Kill residual Chrome processes (Windows example) system("taskkill /im chrome.exe /f", ignore.stdout = TRUE) # Initialize driver with matching Chrome version and custom port rmDr <- rsDriver(browser = "chrome", chromever = "114.0.5735.90", port = 4567L) myclient <- rmDr$client # Navigate to the Google reviews page myclient$navigate("https://www.google.co.uk/search?q=queen%27s+hospital+romford&oq=queen%27s+hospitql+&aqs=chrome.1.69i57j0l5.5843j0j4&sourceid=chrome&ie=UTF-8#lrd=0x47d8a4ce4aaaba81:0xf1185c71ae14d00,1,,,") # Add a small delay to let the page load fully (anti-scraping best practice) Sys.sleep(3) # Get page source pagesource <- myclient$getPageSource()[[1]] # Proceed with your SelectorGadget scraping here... # Don't forget to clean up! myclient$close() rmDr$server$stop()
A quick note: Google has strict anti-scraping measures, so avoid sending too many requests in a short time. Adding delays like Sys.sleep() between actions helps prevent getting blocked.
内容的提问来源于stack exchange,提问作者Varun

