You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Headless Chrome爬取GoodRead页面遇TypeError错误求助

解决Selenium Headless模式爬取Goodreads页面的参数错误问题

你遇到的TypeError: WebDriver.__init__() got an unexpected keyword argument 'chrome_options',是因为新版Selenium(4.x及以上版本)已弃用chrome_options参数,改用options传递浏览器配置,同时需要通过ChromeService管理驱动路径。

以下是修正后的完整无浏览器爬取脚本:

from bs4 import BeautifulSoup
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service as ChromeService
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 配置Chrome Headless模式
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")  # 新版无头模式,兼容性更强
options.add_argument("--disable-gpu")  # 禁用GPU加速,避免环境兼容问题
options.add_argument("--window-size=1920,1080")  # 指定窗口尺寸,保证页面元素布局正常

# 初始化Chrome驱动(符合Selenium 4.x API规范)
driver = webdriver.Chrome(
    service=ChromeService(ChromeDriverManager().install()),
    options=options
)

url = "https://www.goodreads.com/book/show/20656089-indian-polity"

# 加载目标页面
driver.get(url)

# 等待"更多"按钮可点击并触发点击
wait = WebDriverWait(driver, 10)
more_button = wait.until(EC.element_to_be_clickable((By.XPATH, '//span[contains(text(), "...more")]/parent::a')))
more_button.click()

# 获取点击后的页面源码并关闭驱动
page_content = driver.page_source
driver.quit()

# 解析页面数据
soup = BeautifulSoup(page_content, 'html.parser')

# 提取封面图片URL
book_image_element = soup.find("img", {"class": "ResponsiveImage"})
book_image_url = book_image_element.get("src") if book_image_element else "Image not available"

# 提取书籍描述
description_text = "Description element not found in the HTML."
description_element = soup.find("div", {"data-testid": "description"})
if description_element:
    span_element = description_element.find("div", class_="DetailsLayoutRightParagraph").find("span", class_="Formatted")
    description_text = span_element.get_text() if span_element else "Description not available within the expected structure."

# 提取分类标签
genre_names = []
genres_element = soup.find("div", {"data-testid": "genresList"})
if genres_element:
    genre_buttons = genres_element.find_all("span", class_="Button__labelItem")
    genre_names = [genre.get_text() for genre in genre_buttons]

# 保存数据到Excel
data = {
    'Book Image URL': [book_image_url],
    'Description': [description_text],
    'Genre': [', '.join(genre_names)]
}
df = pd.DataFrame(data)
excel_filename = 'book_data.xlsx'
df.to_excel(excel_filename, index=False)

print(f"Data saved to {excel_filename}")

关键修改说明

  • 参数规范修正:将原代码中的chrome_options=options改为options=options,同时通过service参数传入ChromeService实例,适配Selenium 4.x的API要求。
  • Headless模式优化:使用--headless=new(Selenium 4.8+支持的新版无头模式,行为更接近正常浏览器),添加--disable-gpu和窗口尺寸参数,避免页面渲染异常。
  • 逻辑整合:合并原脚本中重复的页面请求操作,直接使用Selenium加载后的页面源码完成所有数据提取,减少冗余步骤。

内容的提问来源于stack exchange,提问作者Abid Fakhrealam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 09:43:38