Python Selenium ChromeDriver报错:WebDriver.__init__()意外参数executable_path
修复Selenium驱动初始化错误及Yelp评论爬取分析建议
一、修复executable_path参数错误
Selenium 4.x版本起,移除了executable_path这个初始化参数,需通过Service类指定ChromeDriver路径。具体修改如下:
1. 新增导入语句
在原有导入代码中添加Service类:
from selenium.webdriver.chrome.service import Service
2. 修改驱动初始化代码
替换原有的驱动创建代码:
# 原代码 driver = webdriver.Chrome(executable_path=chromedriver_path, options=chrome_options) # 修改后代码 service = Service(chromedriver_path) driver = webdriver.Chrome(service=service, options=chrome_options)
修复后的完整代码(含元素选择器适配Yelp最新结构)
!pip install selenium beautifulsoup4 pandas from selenium import webdriver from selenium.webdriver.chrome.service import Service from bs4 import BeautifulSoup import pandas as pd import os import time import random # Set the path to the ChromeDriver executable chromedriver_path = "C:\\Users\\5mxz2\\Downloads\\chromedriver\\chromedriver" # Set the URL of the Yelp page you want to scrape url = "https://www.yelp.com/biz/gelati-celesti-virginia-beach-2" # Set the options for Chrome chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--headless") # 添加反爬参数 chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") chrome_options.add_argument("--disable-blink-features=AutomationControlled") # Create the ChromeDriver instance service = Service(chromedriver_path) driver = webdriver.Chrome(service=service, options=chrome_options) # Load the Yelp page driver.get(url) # 随机等待页面加载 time.sleep(random.uniform(2, 4)) # Extract the page source and pass it to BeautifulSoup soup = BeautifulSoup(driver.page_source, "html.parser") # Find all review elements on the page(适配Yelp最新评论容器class) reviews = soup.find_all("div", class_="review__09f24__oHr9V") # Create empty lists to store the extracted data review_texts = [] ratings = [] dates = [] # Iterate over each review element for review in reviews: # Extract the review text try: review_text = review.find("p", class_="comment__09f24__gu0rG").get_text().strip() except AttributeError: review_text = None review_texts.append(review_text) # Extract the rating try: rating = review.find("div", class_="i-stars__09f24__M1AR7").get("aria-label") except AttributeError: rating = None ratings.append(rating) # Extract the date try: date = review.find("span", class_="css-chan6m").get_text().strip() except AttributeError: date = None dates.append(date) # Create a DataFrame from the extracted data data = { "Review Text": review_texts, "Rating": ratings, "Date": dates } df = pd.DataFrame(data) # Print the DataFrame print(df) # Save the DataFrame as a CSV file csv_path = os.path.join(os.getcwd(), "yelp_reviews.csv") df.to_csv(csv_path, index=False) # Close the ChromeDriver instance driver.quit()
二、Yelp评论爬取与分析实用建议
爬取相关建议
- 反爬规避:
- 除代码中已添加的UA和反检测参数,还可添加
--no-sandbox、--disable-dev-shm-usage参数提升稳定性; - 不要连续爬取大量页面,每爬1-2页添加3-5秒的随机等待;
- 若需爬取多个商家,可切换IP或使用代理池,避免IP被封禁。
- 除代码中已添加的UA和反检测参数,还可添加
- 分页处理:当前代码仅爬第一页,可通过定位"Next"按钮(通常class含
next-link),循环点击直到按钮不可用,实现全量评论爬取; - 元素选择器维护:Yelp会频繁更新页面class,若出现数据提取为空的情况,需打开浏览器开发者工具重新定位元素。
Pandas分析建议
- 数据清洗:将评分文本转换为数字,方便统计:
df["Rating"] = df["Rating"].str.extract(r'(\d+)').astype(float) - 基础统计:用
df["Rating"].describe()查看评分分布,df["Rating"].value_counts(normalize=True)计算各评分占比; - 文本分析:统计评论字数分布,或用TextBlob做简单情感分析:
from textblob import TextBlob df["Sentiment"] = df["Review Text"].apply(lambda x: TextBlob(str(x)).sentiment.polarity if x else None) - 时间趋势分析:将日期转为datetime格式,按月份分组查看评分变化:
df["Date"] = pd.to_datetime(df["Date"], format="%m/%d/%Y") monthly_rating = df.groupby(df["Date"].dt.to_period("M"))["Rating"].mean() monthly_rating.plot(kind="line", title="Monthly Average Rating")
内容的提问来源于stack exchange,提问作者Y0hno
相关产品推荐
相关产品推荐

