You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows环境下Python Selenium Chrome二进制文件路径错误求助

Chrome二进制文件路径错误排查(Yelp评论抓取场景)

我是编程新手,想抓取Yelp平台的评论数据并用Pandas分析,尝试用Selenium+BeautifulSoup实现自动化抓取,但一直卡在Chrome二进制文件路径错误上。以下是我的代码和报错信息:

!pip install selenium
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup
import pandas as pd
import os

# Set the path to the ChromeDriver executable
chromedriver_path = "C:\\Users\\5mxz2\\Downloads\\chromedriver_win32\\chromedriver"

# Set the path to the Chrome binary
chrome_binary_path = "C:\\Program Files\\Google\\Chrome\\Application\\chrome.exe"  # Update this with the correct path to your Chrome binary

# Set the URL of the Yelp page you want to scrape
url = "https://www.yelp.com/biz/gelati-celesti-virginia-beach-2"

# Set the options for Chrome
chrome_options = webdriver.ChromeOptions()
chrome_options.add_argument("--headless")  # Run Chrome in headless mode, comment this line if you want to see the browser window
chrome_options.binary_location = chrome_binary_path

# Create the ChromeDriver service
service = Service(chromedriver_path)

# Create the ChromeDriver instance
driver = webdriver.Chrome(service=service, options=chrome_options)

# Load the Yelp page
driver.get(url)

# Extract the page source and pass it to BeautifulSoup
soup = BeautifulSoup(driver.page_source, "html.parser")

# Find all review elements on the page
reviews = soup.find_all("div", class_="review")

# Create empty lists to store the extracted data
review_texts = []
ratings = []
dates = []

# Iterate over each review element
for review in reviews:
    # Extract the review text
    review_text = review.find("p", class_="comment").get_text()
    review_texts.append(review_text.strip())

    # Extract the rating
    rating = review.find("div", class_="rating").get("aria-label")
    ratings.append(rating)

    # Extract the date
    date = review.find("span", class_="rating-qualifier").get_text()
    dates.append(date.strip())

# Create a DataFrame from the extracted data
data = {
    "Review Text": review_texts,
    "Rating": ratings,
    "Date": dates
}
df = pd.DataFrame(data)

# Print the DataFrame
print(df)

# Get the current working directory
path = os.getcwd()

# Save the DataFrame as a CSV file
csv_path = os.path.join(path, "yelp_reviews.csv")
df.to_csv(csv_path, index=False)

# Close the ChromeDriver instance
driver.quit()

报错信息:

WebDriverException                        Traceback (most recent call last)
<ipython-input-11-6c92e956c704> in <cell line: 27>()
     25 
     26 # Create the ChromeDriver instance
---> 27 driver = webdriver.Chrome(service=service, options=chrome_options)
     28 
     29 # Load the Yelp page

5 frames
/usr/local/lib/python3.10/dist-packages/selenium/webdriver/remote/errorhandler.py in check_response(self, response)
    243                 alert_text = value["alert"].get("text")
    244             raise exception_class(message, screen, stacktrace, alert_text)  # type: ignore[call-arg]  # mypy is not smart enough here
---> 245         raise exception_class(message, screen, stacktrace)

WebDriverException: Message: unknown error: no chrome binary at C:\Program Files\Google\Chrome\Application\chrome.exe
Stacktrace:
#0 0x55f8912b24e3 <unknown>
#1 0x55f890fe1c76 <unknown>
#2 0x55f8910085e0 <unknown>
#3 0x55f891007029 <unknown>
#4 0x55f891045ccc <unknown>
#5 0x55f89104547f <unknown>
#6 0x55f89103cde3 <unknown>
#7 0x55f8910122dd <unknown>
#8 0x55f89101334e <unknown>
#9 0x55f8912723e4 <unknown>
#10 0x55f8912763d7 <unknown>
#11 0x55f891280b20 <unknown>
#12 0x55f891277023 <unknown>
#13 0x55f8912451aa <unknown>
#14 0x55f89129b6b8 <unknown>
#15 0x55f89129b847 <unknown>
#16 0x55f8912ab243 <unknown>
#17 0x7ff7aa929609 start_thread

问题根源

报错明确指出找不到指定路径下的Chrome二进制文件,核心原因要么是路径填写错误,要么是你运行代码的环境(比如Colab、Linux服务器)和你写的Windows路径不匹配。

修复步骤

1. 确认Chrome实际安装路径

  • Windows系统:打开Chrome → 右上角三个点 → 帮助 → 关于Google Chrome,查看安装路径。常见路径除了C:\Program Files\Google\Chrome\Application\chrome.exe,还可能是C:\Program Files (x86)\Google\Chrome\Application\chrome.exe(32位版本)。
  • Linux系统:Chrome二进制路径一般是/usr/bin/google-chrome。
  • Mac系统:路径通常是/Applications/Google Chrome.app/Contents/MacOS/Google Chrome。

2. 匹配运行环境的路径

从报错里的/usr/local/lib/python3.10/dist-packages来看,你大概率是在Linux环境(比如Colab)运行代码,但写了Windows路径,直接替换成Linux对应的路径即可。

3. 可选:简化配置

如果Chrome是默认安装,Selenium会自动查找系统里的Chrome,直接注释掉手动设置二进制路径的代码:

# chrome_options.binary_location = chrome_binary_path

4. 验证ChromeDriver与Chrome版本匹配

确保ChromeDriver的版本和你安装的Chrome版本完全一致,版本不匹配也会导致各类错误。可以在Chrome“关于”页面看版本,再下载对应版本的ChromeDriver。

修复后的示例代码(Linux/Colab环境)

!pip install selenium beautifulsoup4 pandas
!apt install chromium-chromedriver  # Colab环境需要安装ChromeDriver
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from bs4 import BeautifulSoup
import pandas as pd
import os

# Linux/Colab环境下的路径
chromedriver_path = "/usr/bin/chromedriver"
chrome_binary_path = "/usr/bin/google-chrome"

url = "https://www.yelp.com/biz/gelati-celesti-virginia-beach-2"

chrome_options = webdriver.ChromeOptions()
chrome_options.add_argument("--headless")
chrome_options.add_argument("--no-sandbox")  # Linux环境必填参数
chrome_options.add_argument("--disable-dev-shm-usage")  # 解决内存不足问题
chrome_options.binary_location = chrome_binary_path

service = Service(chromedriver_path)
driver = webdriver.Chrome(service=service, options=chrome_options)

driver.get(url)
soup = BeautifulSoup(driver.page_source, "html.parser")

# 注意:Yelp页面结构可能更新,这里适配了最新的选择器
reviews = soup.find_all("div", class_="review__09f24__oHr9V")

review_texts = []
ratings = []
dates = []

for review in reviews:
    # 提取评论文本
    review_text_elem = review.find("p", class_="comment__09f24__gu0rG")
    review_texts.append(review_text_elem.get_text().strip() if review_text_elem else None)
    
    # 提取评分
    rating_elem = review.find("div", class_="rating__09f24__mLAFf")
    ratings.append(rating_elem.get("aria-label") if rating_elem else None)
    
    # 提取日期
    date_elem = review.find("span", class_="css-chan6m")
    dates.append(date_elem.get_text().strip() if date_elem else None)

data = {
    "Review Text": review_texts,
    "Rating": ratings,
    "Date": dates
}
df = pd.DataFrame(data)

print(df)

csv_path = os.path.join(os.getcwd(), "yelp_reviews.csv")
df.to_csv(csv_path, index=False)

driver.quit()

内容的提问来源于stack exchange,提问作者Y0hno

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 10:32:34