Python2.7用Selenium爬Google Trends遇options参数TypeError求助
问题描述
我目前使用以下Python代码,试图通过Selenium和BeautifulSoup4提取Google Trends的博客文章标题:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.common.action_chains import ActionChains from bs4 import BeautifulSoup chrome_options = Options() chrome_options.add_argument('--headless') # Run Chrome in headless mode chrome_options.add_argument('--no-sandbox') # Add this option to avoid sandbox issues driver = webdriver.Chrome(options=chrome_options) driver.maximize_window() try: driver.get('https://trends.google.com/trends/trendingsearches/realtime?geo=US&hl=en-US&category=t') scrollable_element = driver.find_element(By.XPATH, '//body') actions = ActionChains(driver) actions.move_to_element(scrollable_element) actions.perform() page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') title_element = soup.find('div', class_='summary-text') title = title_element.text.strip() if title_element else None print("Title:", title) except Exception as e: print("An error occurred:", str(e)) finally: driver.quit()
但运行后返回如下错误:
Traceback (most recent call last): File "extract_titles.py", line 11, in <module> driver = webdriver.Chrome(options=chrome_options) TypeError: __init__() got an unexpected keyword argument 'options'
我的限制条件:
- 必须使用Python 2.7
- 服务器无Node.js环境,无法使用Puppeteer
- 尝试过PHP Simple HTML DOM和DOMDocument,但无法处理Google Trends的动态内容
- 选择使用Selenium解决动态内容问题
请问如何修复该错误并获取博客文章标题?或能否提供使用其他工具的Python代码方案?
解决方案
1. 修复Selenium初始化错误
Python 2.7对应的Selenium版本(多为Selenium 3.x及以下)中,webdriver.Chrome初始化不支持options参数,需改用chrome_options参数。修改初始化代码:
# 替换原代码中的driver初始化行 driver = webdriver.Chrome(chrome_options=chrome_options)
若出现By类找不到的错误,可确认导入路径,旧版Selenium中该类路径通常不变,若仍报错可直接改用元素定位的旧方法(如find_element_by_xpath)替代By.XPATH。
2. 优化内容提取逻辑(修复后完整代码)
原代码仅提取单个标题,且未等待页面动态加载完成,容易出现元素找不到的问题。以下是优化后的完整代码:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.common.action_chains import ActionChains from bs4 import BeautifulSoup import time chrome_options = Options() chrome_options.add_argument('--headless') chrome_options.add_argument('--no-sandbox') chrome_options.add_argument('--disable-dev-shm-usage') # 解决Linux服务器内存不足问题 driver = webdriver.Chrome(chrome_options=chrome_options) driver.maximize_window() try: driver.get('https://trends.google.com/trends/trendingsearches/realtime?geo=US&hl=en-US&category=t') # 等待页面动态内容渲染 time.sleep(3) # 滚动页面加载更多内容 scrollable_element = driver.find_element(By.XPATH, '//body') actions = ActionChains(driver) actions.move_to_element(scrollable_element).perform() time.sleep(2) # 等待滚动后内容加载 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 提取所有博客文章标题 title_elements = soup.find_all('div', class_='summary-text') if title_elements: print("提取到的标题:") for idx, elem in enumerate(title_elements, 1): title = elem.text.strip() print(f"{idx}. {title}") else: print("未找到任何标题元素") except Exception as e: print("发生错误:", str(e)) finally: driver.quit()
3. 替代方案:使用requests-html(支持Python 2.7)
若Selenium配置繁琐,可尝试requests-html,它内置JavaScript渲染能力,无需浏览器驱动。
首先安装适配Python 2.7的版本:
pip install requests-html==0.10.0
然后编写提取代码:
from requests_html import HTMLSession session = HTMLSession() try: r = session.get('https://trends.google.com/trends/trendingsearches/realtime?geo=US&hl=en-US&category=t') # 渲染JavaScript内容,滚动2次加载更多 r.html.render(sleep=3, scrolldown=2) # 提取所有标题 titles = r.html.find('.summary-text') if titles: print("提取到的标题:") for idx, title in enumerate(titles, 1): print(f"{idx}. {title.text.strip()}") else: print("未找到标题") except Exception as e: print("发生错误:", str(e)) finally: session.close()
内容的提问来源于stack exchange,提问作者Hassan Suriya
相关产品推荐
相关产品推荐

