You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python2.7用Selenium爬Google Trends遇options参数TypeError求助

问题描述

我目前使用以下Python代码,试图通过Selenium和BeautifulSoup4提取Google Trends的博客文章标题:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.common.action_chains import ActionChains
from bs4 import BeautifulSoup

chrome_options = Options()
chrome_options.add_argument('--headless')  # Run Chrome in headless mode
chrome_options.add_argument('--no-sandbox')  # Add this option to avoid sandbox issues
driver = webdriver.Chrome(options=chrome_options)
driver.maximize_window()
try:
    driver.get('https://trends.google.com/trends/trendingsearches/realtime?geo=US&hl=en-US&category=t')
    scrollable_element = driver.find_element(By.XPATH, '//body')
    actions = ActionChains(driver)
    actions.move_to_element(scrollable_element)
    actions.perform()

    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'html.parser')
    title_element = soup.find('div', class_='summary-text')
    title = title_element.text.strip() if title_element else None
    print("Title:", title)

except Exception as e:
    print("An error occurred:", str(e))

finally:
    driver.quit()

但运行后返回如下错误:

Traceback (most recent call last):
  File "extract_titles.py", line 11, in <module>
    driver = webdriver.Chrome(options=chrome_options)
TypeError: __init__() got an unexpected keyword argument 'options'

我的限制条件:

  • 必须使用Python 2.7
  • 服务器无Node.js环境,无法使用Puppeteer
  • 尝试过PHP Simple HTML DOM和DOMDocument,但无法处理Google Trends的动态内容
  • 选择使用Selenium解决动态内容问题

请问如何修复该错误并获取博客文章标题?或能否提供使用其他工具的Python代码方案?

解决方案

1. 修复Selenium初始化错误

Python 2.7对应的Selenium版本(多为Selenium 3.x及以下)中,webdriver.Chrome初始化不支持options参数,需改用chrome_options参数。修改初始化代码:

# 替换原代码中的driver初始化行
driver = webdriver.Chrome(chrome_options=chrome_options)

若出现By类找不到的错误,可确认导入路径,旧版Selenium中该类路径通常不变,若仍报错可直接改用元素定位的旧方法(如find_element_by_xpath)替代By.XPATH。

2. 优化内容提取逻辑(修复后完整代码)

原代码仅提取单个标题,且未等待页面动态加载完成,容易出现元素找不到的问题。以下是优化后的完整代码:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.common.action_chains import ActionChains
from bs4 import BeautifulSoup
import time

chrome_options = Options()
chrome_options.add_argument('--headless')
chrome_options.add_argument('--no-sandbox')
chrome_options.add_argument('--disable-dev-shm-usage')  # 解决Linux服务器内存不足问题
driver = webdriver.Chrome(chrome_options=chrome_options)
driver.maximize_window()
try:
    driver.get('https://trends.google.com/trends/trendingsearches/realtime?geo=US&hl=en-US&category=t')
    # 等待页面动态内容渲染
    time.sleep(3)
    # 滚动页面加载更多内容
    scrollable_element = driver.find_element(By.XPATH, '//body')
    actions = ActionChains(driver)
    actions.move_to_element(scrollable_element).perform()
    time.sleep(2)  # 等待滚动后内容加载

    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'html.parser')
    # 提取所有博客文章标题
    title_elements = soup.find_all('div', class_='summary-text')
    if title_elements:
        print("提取到的标题:")
        for idx, elem in enumerate(title_elements, 1):
            title = elem.text.strip()
            print(f"{idx}. {title}")
    else:
        print("未找到任何标题元素")

except Exception as e:
    print("发生错误:", str(e))

finally:
    driver.quit()

3. 替代方案:使用requests-html(支持Python 2.7)

若Selenium配置繁琐,可尝试requests-html,它内置JavaScript渲染能力,无需浏览器驱动。

首先安装适配Python 2.7的版本:

pip install requests-html==0.10.0

然后编写提取代码:

from requests_html import HTMLSession

session = HTMLSession()
try:
    r = session.get('https://trends.google.com/trends/trendingsearches/realtime?geo=US&hl=en-US&category=t')
    # 渲染JavaScript内容,滚动2次加载更多
    r.html.render(sleep=3, scrolldown=2)
    # 提取所有标题
    titles = r.html.find('.summary-text')
    if titles:
        print("提取到的标题:")
        for idx, title in enumerate(titles, 1):
            print(f"{idx}. {title.text.strip()}")
    else:
        print("未找到标题")
except Exception as e:
    print("发生错误:", str(e))
finally:
    session.close()

内容的提问来源于stack exchange,提问作者Hassan Suriya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 02:32:47