Windows环境下Python Selenium Chrome二进制文件路径错误求助
Chrome二进制文件路径错误排查(Yelp评论抓取场景)
我是编程新手,想抓取Yelp平台的评论数据并用Pandas分析,尝试用Selenium+BeautifulSoup实现自动化抓取,但一直卡在Chrome二进制文件路径错误上。以下是我的代码和报错信息:
!pip install selenium from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from bs4 import BeautifulSoup import pandas as pd import os # Set the path to the ChromeDriver executable chromedriver_path = "C:\\Users\\5mxz2\\Downloads\\chromedriver_win32\\chromedriver" # Set the path to the Chrome binary chrome_binary_path = "C:\\Program Files\\Google\\Chrome\\Application\\chrome.exe" # Update this with the correct path to your Chrome binary # Set the URL of the Yelp page you want to scrape url = "https://www.yelp.com/biz/gelati-celesti-virginia-beach-2" # Set the options for Chrome chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--headless") # Run Chrome in headless mode, comment this line if you want to see the browser window chrome_options.binary_location = chrome_binary_path # Create the ChromeDriver service service = Service(chromedriver_path) # Create the ChromeDriver instance driver = webdriver.Chrome(service=service, options=chrome_options) # Load the Yelp page driver.get(url) # Extract the page source and pass it to BeautifulSoup soup = BeautifulSoup(driver.page_source, "html.parser") # Find all review elements on the page reviews = soup.find_all("div", class_="review") # Create empty lists to store the extracted data review_texts = [] ratings = [] dates = [] # Iterate over each review element for review in reviews: # Extract the review text review_text = review.find("p", class_="comment").get_text() review_texts.append(review_text.strip()) # Extract the rating rating = review.find("div", class_="rating").get("aria-label") ratings.append(rating) # Extract the date date = review.find("span", class_="rating-qualifier").get_text() dates.append(date.strip()) # Create a DataFrame from the extracted data data = { "Review Text": review_texts, "Rating": ratings, "Date": dates } df = pd.DataFrame(data) # Print the DataFrame print(df) # Get the current working directory path = os.getcwd() # Save the DataFrame as a CSV file csv_path = os.path.join(path, "yelp_reviews.csv") df.to_csv(csv_path, index=False) # Close the ChromeDriver instance driver.quit()
报错信息:
WebDriverException Traceback (most recent call last) <ipython-input-11-6c92e956c704> in <cell line: 27>() 25 26 # Create the ChromeDriver instance ---> 27 driver = webdriver.Chrome(service=service, options=chrome_options) 28 29 # Load the Yelp page 5 frames /usr/local/lib/python3.10/dist-packages/selenium/webdriver/remote/errorhandler.py in check_response(self, response) 243 alert_text = value["alert"].get("text") 244 raise exception_class(message, screen, stacktrace, alert_text) # type: ignore[call-arg] # mypy is not smart enough here ---> 245 raise exception_class(message, screen, stacktrace) WebDriverException: Message: unknown error: no chrome binary at C:\Program Files\Google\Chrome\Application\chrome.exe Stacktrace: #0 0x55f8912b24e3 <unknown> #1 0x55f890fe1c76 <unknown> #2 0x55f8910085e0 <unknown> #3 0x55f891007029 <unknown> #4 0x55f891045ccc <unknown> #5 0x55f89104547f <unknown> #6 0x55f89103cde3 <unknown> #7 0x55f8910122dd <unknown> #8 0x55f89101334e <unknown> #9 0x55f8912723e4 <unknown> #10 0x55f8912763d7 <unknown> #11 0x55f891280b20 <unknown> #12 0x55f891277023 <unknown> #13 0x55f8912451aa <unknown> #14 0x55f89129b6b8 <unknown> #15 0x55f89129b847 <unknown> #16 0x55f8912ab243 <unknown> #17 0x7ff7aa929609 start_thread
问题根源
报错明确指出找不到指定路径下的Chrome二进制文件,核心原因要么是路径填写错误,要么是你运行代码的环境(比如Colab、Linux服务器)和你写的Windows路径不匹配。
修复步骤
1. 确认Chrome实际安装路径
- Windows系统:打开Chrome → 右上角三个点 → 帮助 → 关于Google Chrome,查看安装路径。常见路径除了
C:\Program Files\Google\Chrome\Application\chrome.exe,还可能是C:\Program Files (x86)\Google\Chrome\Application\chrome.exe(32位版本)。 - Linux系统:Chrome二进制路径一般是
/usr/bin/google-chrome。 - Mac系统:路径通常是
/Applications/Google Chrome.app/Contents/MacOS/Google Chrome。
2. 匹配运行环境的路径
从报错里的/usr/local/lib/python3.10/dist-packages来看,你大概率是在Linux环境(比如Colab)运行代码,但写了Windows路径,直接替换成Linux对应的路径即可。
3. 可选:简化配置
如果Chrome是默认安装,Selenium会自动查找系统里的Chrome,直接注释掉手动设置二进制路径的代码:
# chrome_options.binary_location = chrome_binary_path
4. 验证ChromeDriver与Chrome版本匹配
确保ChromeDriver的版本和你安装的Chrome版本完全一致,版本不匹配也会导致各类错误。可以在Chrome“关于”页面看版本,再下载对应版本的ChromeDriver。
修复后的示例代码(Linux/Colab环境)
!pip install selenium beautifulsoup4 pandas !apt install chromium-chromedriver # Colab环境需要安装ChromeDriver from selenium import webdriver from selenium.webdriver.chrome.service import Service from bs4 import BeautifulSoup import pandas as pd import os # Linux/Colab环境下的路径 chromedriver_path = "/usr/bin/chromedriver" chrome_binary_path = "/usr/bin/google-chrome" url = "https://www.yelp.com/biz/gelati-celesti-virginia-beach-2" chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--headless") chrome_options.add_argument("--no-sandbox") # Linux环境必填参数 chrome_options.add_argument("--disable-dev-shm-usage") # 解决内存不足问题 chrome_options.binary_location = chrome_binary_path service = Service(chromedriver_path) driver = webdriver.Chrome(service=service, options=chrome_options) driver.get(url) soup = BeautifulSoup(driver.page_source, "html.parser") # 注意:Yelp页面结构可能更新,这里适配了最新的选择器 reviews = soup.find_all("div", class_="review__09f24__oHr9V") review_texts = [] ratings = [] dates = [] for review in reviews: # 提取评论文本 review_text_elem = review.find("p", class_="comment__09f24__gu0rG") review_texts.append(review_text_elem.get_text().strip() if review_text_elem else None) # 提取评分 rating_elem = review.find("div", class_="rating__09f24__mLAFf") ratings.append(rating_elem.get("aria-label") if rating_elem else None) # 提取日期 date_elem = review.find("span", class_="css-chan6m") dates.append(date_elem.get_text().strip() if date_elem else None) data = { "Review Text": review_texts, "Rating": ratings, "Date": dates } df = pd.DataFrame(data) print(df) csv_path = os.path.join(os.getcwd(), "yelp_reviews.csv") df.to_csv(csv_path, index=False) driver.quit()
内容的提问来源于stack exchange,提问作者Y0hno
相关产品推荐
相关产品推荐

