使用Headless Chrome爬取GoodRead页面遇TypeError错误求助
解决Selenium Headless模式爬取Goodreads页面的参数错误问题
你遇到的TypeError: WebDriver.__init__() got an unexpected keyword argument 'chrome_options',是因为新版Selenium(4.x及以上版本)已弃用chrome_options参数,改用options传递浏览器配置,同时需要通过ChromeService管理驱动路径。
以下是修正后的完整无浏览器爬取脚本:
from bs4 import BeautifulSoup import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service as ChromeService from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 配置Chrome Headless模式 options = webdriver.ChromeOptions() options.add_argument("--headless=new") # 新版无头模式,兼容性更强 options.add_argument("--disable-gpu") # 禁用GPU加速,避免环境兼容问题 options.add_argument("--window-size=1920,1080") # 指定窗口尺寸,保证页面元素布局正常 # 初始化Chrome驱动(符合Selenium 4.x API规范) driver = webdriver.Chrome( service=ChromeService(ChromeDriverManager().install()), options=options ) url = "https://www.goodreads.com/book/show/20656089-indian-polity" # 加载目标页面 driver.get(url) # 等待"更多"按钮可点击并触发点击 wait = WebDriverWait(driver, 10) more_button = wait.until(EC.element_to_be_clickable((By.XPATH, '//span[contains(text(), "...more")]/parent::a'))) more_button.click() # 获取点击后的页面源码并关闭驱动 page_content = driver.page_source driver.quit() # 解析页面数据 soup = BeautifulSoup(page_content, 'html.parser') # 提取封面图片URL book_image_element = soup.find("img", {"class": "ResponsiveImage"}) book_image_url = book_image_element.get("src") if book_image_element else "Image not available" # 提取书籍描述 description_text = "Description element not found in the HTML." description_element = soup.find("div", {"data-testid": "description"}) if description_element: span_element = description_element.find("div", class_="DetailsLayoutRightParagraph").find("span", class_="Formatted") description_text = span_element.get_text() if span_element else "Description not available within the expected structure." # 提取分类标签 genre_names = [] genres_element = soup.find("div", {"data-testid": "genresList"}) if genres_element: genre_buttons = genres_element.find_all("span", class_="Button__labelItem") genre_names = [genre.get_text() for genre in genre_buttons] # 保存数据到Excel data = { 'Book Image URL': [book_image_url], 'Description': [description_text], 'Genre': [', '.join(genre_names)] } df = pd.DataFrame(data) excel_filename = 'book_data.xlsx' df.to_excel(excel_filename, index=False) print(f"Data saved to {excel_filename}")
关键修改说明
- 参数规范修正:将原代码中的
chrome_options=options改为options=options,同时通过service参数传入ChromeService实例,适配Selenium 4.x的API要求。 - Headless模式优化:使用
--headless=new(Selenium 4.8+支持的新版无头模式,行为更接近正常浏览器),添加--disable-gpu和窗口尺寸参数,避免页面渲染异常。 - 逻辑整合:合并原脚本中重复的页面请求操作,直接使用Selenium加载后的页面源码完成所有数据提取,减少冗余步骤。
内容的提问来源于stack exchange,提问作者Abid Fakhrealam
相关产品推荐
相关产品推荐

